Component

Filing extraction

The published reporting tables pulled, checked against the agency's own arithmetic, and delivered as files on a schedule. Every figure carries the file and the cell it was read from.

What it is

An agency publishes a table. It is a spreadsheet with merged headers, or a PDF, or a portal that returns one system at a time. We pull it, read it, check the totals against the components the same file publishes, and hand you the result as data.

The checking is most of the value. A published total and the components it is made of can disagree, and until somebody adds the components up nothing says so.

What you hand us

The publication, or a link to it. If it is behind a login or a request process, the access. We do not scrape around a gate.

What you get back

One row per record, with the figures as published, the figures as computed from their own components where those differ, and the coordinate each was read from. Re-run on your schedule.

What it is built from

Your source, plus a map recording which header each field came from and what was done to it. That map is what lets a figure be traced back to a cell instead of to a process.

What it does not do

It does not correct the source. Where a published figure is wrong we report the discrepancy and publish both numbers; we do not substitute our own and we do not quietly repair a total.

It does not read a scanned page. A PDF of an image needs a person, and we will say so before quoting rather than after.

It does not give you an API. You get files on a schedule; there is no endpoint to integrate against and no account to set up.

Start a conversation

A short call is usually enough to say whether we are the right people for it.

Email ammar@municorn.us

Related

Cite as

Municorn.us, “Filing extraction”, https://municorn.us/build/filing-extraction/ (updated 2026-09-08).

CC BY 4.0 — reuse and redistribution are permitted with attribution to Municorn.us and a link to this page. Terms.