Decision Brief · Prepared for review
A national civic-record contribution, waiting on decisions only you can make.
Ten vendors host the agendas and minutes of most US local government. I've already pulled four corpora to disk — 1.4 million documents, 18 million pages, 2.18 TB — one vendor fully reconciled as proof of method; the rest are mechanically reachable. Nothing here is blocked on engineering — only on scope and convention.
What exists on disk today — measured 2026-08-08
| corpus | PDF files | pages | size | status |
|---|---|---|---|---|
| Legistar | 417,636 | 2,769,809 | 140.9 GB | reconciled |
| CivicClerk | 140,543 | 2,431,717 | 303.3 GB | collected |
| CivicPlus | 71,664 | 2,165,944 | 237.1 GB | collected |
| current subtotal | 629,843 | 7,367,470 | 681.3 GB | |
| CivicPlus 2022 | 789,687 | 10,691,423 | 1.5 TB | legacy · assessment |
| total | 1,419,530 | 18,058,893 | 2.18 TB |
Legistar is the fully-reconciled proof of method — 170 jurisdictions across 33 states, every file traced to a source URL and event id, gaps at 0.48% and each individually classified. CivicClerk and CivicPlus are collected and still being reconciled; the 2022 CivicPlus set is the legacy corpus under forensic assessment — and it holds nearly all of the ~20,000 unreadable files, plus a further ~13,550 that are valid but blank (empty print-to-PDF shells that render to nothing), because decay concentrates in the oldest data. Full build-out across all ten vendors is plausibly several million more documents.
The decisions you own — in critical-path order
Collection name Blocking
Trivial to decide, but identifiers cannot be finalized without it. It gates everything downstream.
Entity scope Sizes the corpus
Cities only, or the full civic superset — school boards, library boards, counties, special districts? Recommendation: superset. Same scrapers, same document types, and 4,403 tenants are K-12 school boards from a single platform — among the most contested, least-preserved records in the country.
Delivery model May need hardware
ia CLI upload (ready now) versus a physical drive handoff with data and tooling for IA to ingest. At multi-million-document scale, the path should be chosen before the corpus grows.
Then, and permanent once uploaded: item granularity, and a vendor-neutral identifier scheme keyed on jurisdiction (state · place · Census GEOID). Tenants migrate between vendors — one place has changed vendor four times — and a vendor-prefixed identifier would fracture that history exactly where documents tend to go missing.
Two risks you should hear unprompted
Privacy · confirm with IA
Personal data in agenda packets
Policy is deliberately narrow — agendas, minutes, action summaries only. But where a city publishes the full backup packet at the agenda link, ~2,900 packets (~1.1%, median ~130 pages) slipped in. The genuine exposure is names and home addresses in permit and license applications; HR and disciplinary matters are handled in closed session and don't reach the public record absent clerk error. Measured, and strippable before upload. Worth confirming IA's position on personal data in bulk-republished public records.
Data integrity · load-bearing
Index is not inventory
A vendor API listing a document is not proof it exists. The discipline throughout — "we didn't look" is never recorded as "it isn't there" — is why the provenance figure is trustworthy. It is also the exact failure that silently cost tens of thousands of meetings elsewhere before it was caught.
One inconsistency to resolve first
Hardest to reverse
The two briefs in this effort disagree on item granularity — one recommends place-year (~2,600 items/vendor, manageable), the other meeting-level (~200,000 items, the natural citation unit). Because granularity is the single hardest thing to change after upload, align on it before anything ships. It hinges on one question only IA can answer: does search inside a multi-file item resolve to the specific file, or only to the item?
Context — this is not a cold outreach
Already contributing, on a three-database design
Present
What the worker fleet is doing right now — claim, download, upload to IA.
Past
The authoritative mirror of what is preserved on the Archive, with full history.
Future
Every US local government — the map of what the record should contain.
The same effort already runs a fleet of ~10 workers uploading community video to IA's Community Media collections — today ~3.2 million items across ~2,200 collections. The civic-documents work is the "future tense" of that same design: federate, don't merge.
Read next — 15 minutes, not 76 PDFs
OVERVIEW.pdf — “Preserving Community Media: A Three-Part Effort”
Plain-language, deliberately brief, and well made. It lays out the three-database structure above — present / past / future — and how this contribution fits the mission. The single best orientation before any decision.
Open the PDFThe story behind the numbers
If the figures above are the ledger, this is what they mean — ten short stories, then five paragraphs. (Full text: 00.0-stepping-back-from-CMA-operations/project-stories.md)
- The tower of paper. Four scans, testing the agenda-retrieval software on 1,200 places, 18,058,893 pages across 1.4 million PDFs and 2.18 terabytes — the quiet, unglamorous bulk of American local democracy, the meetings almost nobody attended, preserved anyway.
- The past is where the rot lives. Of the whole haul's 20,474 unreadable files, 19,418 sit in the single oldest corpus — the 2022 CivicPlus forensic set. The newer runs are nearly clean. Decay isn't hypothetical; it has an address, and it's the archive you didn't get to first.
- You can read a city's soul in its mean page count. Legistar averages 6.6 pages a document; CivicPlus averages 30.2. One vendor's world is terse indexes and action summaries; the other's is fat agenda packets. Same democracy, two dialects of paperwork.
- The board that never existed. In Texas, 6,613 documents were filed under a body called “General” — because the software defaulted an uncategorized meeting to that word, and 106 towns never changed it. Four separate councils dissolved into one folder. A field can be full and still say nothing.
- The meeting no one wrote down. 11,553 meetings produced zero files and were recorded nowhere but a scrolling log — not a failure, not a match, just gone. The gap you can't count is the one that quietly becomes a statistic.
- Index is not inventory. The Legistar API will happily list a document it then refuses to hand you. The whole method turns on one rule learned the hard way: “we didn't look” must never be filed as “it isn't there.”
- Honesty, counted. 418,871 of 418,871 files traced back to a source URL and event id; the 1,918 that couldn't be retrieved were each given a name and a reason — a 404, a bare filename, a withdrawn doc — instead of being lumped together as loss.
- A map that didn't exist. Enumerating vendor subdomains from the Wayback index turned a working list of 2,414 into 12,362 tenants across ten platforms — a jurisdiction-to-vendor atlas of who runs American civic meetings that, as far as anyone knows, had never been drawn.
- The address in the permit. Buried in the fat packets ride names and home addresses from permit applications — the ethics problem of republishing the public record in bulk, where “technically already public” meets “now trivially searchable.”
- The clerk who knows every voice. The endgame isn't a scraper; it's Archive Corps — a retired city clerk who recognizes every council member by sound, a librarian who builds the subject vocabulary, volunteers doing the enrichment machines can't.
Somewhere in the United States tonight, a city council is meeting in a half-empty room, and a camera nobody is watching is recording it. Multiply that by tens of thousands of towns and forty years, and you have one of the largest, least-studied records of a democracy anywhere — and one of the most fragile, because it lives on volunteer laptops, dying platforms, and vendor servers that forget. This project is the attempt to catch it before it goes. The tally so far is deceptively dull and secretly enormous: eighteen million pages, one-point-four million documents, two-plus terabytes of agendas and minutes — the connective tissue of self-government, gathered so it can't be quietly deleted.
The most haunting number isn't the biggest one. It's that almost every unreadable file in the entire collection — nineteen thousand of them — belongs to the oldest corpus, the 2022 forensic set. The newer scrapes come back clean; the old one comes back with holes. That's the whole argument for the work in a single statistic: digital records don't fail all at once, they fail at the edges first, and every year you wait, the edges move inward. The forensic bundle isn't just data; it's a memento mori for the born-digital record.
Look closer and the corpus starts to gossip. Legistar towns publish lean, six-page index documents; CivicPlus towns publish thirty-page packets stuffed with backup material. The page-count histogram is really an ethnography — you can see which governments treat the agenda as a pointer and which treat it as the whole file, which is why the same tooling had to learn ten different vendor dialects, and why merely finding the towns meant drawing a 12,362-tenant map that didn't previously exist.
But the real story is epistemological, and it's the one worth telling twice. Again and again the project ran into the same trap wearing different masks: a document the API lists but won't deliver; a meeting that generated no files and so was recorded nowhere; six thousand records filed under a board named “General” that turned out not to be a board at all. Each time, the lesson was the same — a populated field is not a usable value, and silence is not absence. The discipline that came out of it is almost moral: count the thing itself, give every gap its own reason code, and reconcile in both directions until zero files are of unknown origin. That's why “418,871 of 418,871” is the proudest number here. It's not that nothing was lost — nearly two thousand documents were — it's that nothing was lost anonymously.
And it all points somewhere human. The endgame is this contribution to Democracy's Library and a three-tense architecture that reads more like a clock than a schema — collector.db for what the fleet is doing now, archive.db for what's already preserved, civic.db for every government that should be represented but isn't yet. The last database is a to-do list for a democracy's memory. Filling it won't be done by scrapers alone; it'll be done by Archive Corps — the retired clerk who knows every voice, the student learning archival standards, the neighbor who can name the people in a 2009 zoning hearing. The stories in these numbers are about decay and discipline, but the one they're all building toward is about stewardship: a community learning to keep its own record, at a scale no one has tried before.