Decision Brief · Prepared for review

A national civic-record contribution, waiting on decisions only you can make.

Ten vendors host the agendas and minutes of most US local government. I've already pulled four corpora to disk — 1.4 million documents, 18 million pages, 2.18 TB — one vendor fully reconciled as proof of method; the rest are mechanically reachable. Nothing here is blocked on engineering — only on scope and convention.

To
Internet Archive · Democracy's Library US
Re
Scoping and sizing a civic-meeting document contribution
From
Community Media Archive — agendas & minutes effort

What exists on disk today — measured 2026-08-08

1,419,530
PDF files collected across four corpora — three current vendors plus a 2022 legacy set
18,058,893
pages — the corpus of local proceedings
2.18 TB
on disk — collected, not yet uploaded; this is the delivery-model question
12,362
vendor tenants mapped across 10 platforms — the reachable target set
corpusPDF filespagessizestatus
Legistar417,6362,769,809140.9 GBreconciled
CivicClerk140,5432,431,717303.3 GBcollected
CivicPlus71,6642,165,944237.1 GBcollected
current subtotal629,8437,367,470681.3 GB
CivicPlus 2022789,68710,691,4231.5 TBlegacy · assessment
total1,419,53018,058,8932.18 TB

Legistar is the fully-reconciled proof of method — 170 jurisdictions across 33 states, every file traced to a source URL and event id, gaps at 0.48% and each individually classified. CivicClerk and CivicPlus are collected and still being reconciled; the 2022 CivicPlus set is the legacy corpus under forensic assessment — and it holds nearly all of the ~20,000 unreadable files, plus a further ~13,550 that are valid but blank (empty print-to-PDF shells that render to nothing), because decay concentrates in the oldest data. Full build-out across all ten vendors is plausibly several million more documents.

The decisions you own — in critical-path order

1

Collection name Blocking

Trivial to decide, but identifiers cannot be finalized without it. It gates everything downstream.

2

Entity scope Sizes the corpus

Cities only, or the full civic superset — school boards, library boards, counties, special districts? Recommendation: superset. Same scrapers, same document types, and 4,403 tenants are K-12 school boards from a single platform — among the most contested, least-preserved records in the country.

3

Delivery model May need hardware

ia CLI upload (ready now) versus a physical drive handoff with data and tooling for IA to ingest. At multi-million-document scale, the path should be chosen before the corpus grows.

Then, and permanent once uploaded: item granularity, and a vendor-neutral identifier scheme keyed on jurisdiction (state · place · Census GEOID). Tenants migrate between vendors — one place has changed vendor four times — and a vendor-prefixed identifier would fracture that history exactly where documents tend to go missing.

Two risks you should hear unprompted

Privacy · confirm with IA

Personal data in agenda packets

Policy is deliberately narrow — agendas, minutes, action summaries only. But where a city publishes the full backup packet at the agenda link, ~2,900 packets (~1.1%, median ~130 pages) slipped in. The genuine exposure is names and home addresses in permit and license applications; HR and disciplinary matters are handled in closed session and don't reach the public record absent clerk error. Measured, and strippable before upload. Worth confirming IA's position on personal data in bulk-republished public records.

Data integrity · load-bearing

Index is not inventory

A vendor API listing a document is not proof it exists. The discipline throughout — "we didn't look" is never recorded as "it isn't there" — is why the provenance figure is trustworthy. It is also the exact failure that silently cost tens of thousands of meetings elsewhere before it was caught.

One inconsistency to resolve first

Hardest to reverse

The two briefs in this effort disagree on item granularity — one recommends place-year (~2,600 items/vendor, manageable), the other meeting-level (~200,000 items, the natural citation unit). Because granularity is the single hardest thing to change after upload, align on it before anything ships. It hinges on one question only IA can answer: does search inside a multi-file item resolve to the specific file, or only to the item?

Context — this is not a cold outreach

Already contributing, on a three-database design

Present

collector.db

What the worker fleet is doing right now — claim, download, upload to IA.

Past

archive.db

The authoritative mirror of what is preserved on the Archive, with full history.

Future

civic.db

Every US local government — the map of what the record should contain.

The same effort already runs a fleet of ~10 workers uploading community video to IA's Community Media collections — today ~3.2 million items across ~2,200 collections. The civic-documents work is the "future tense" of that same design: federate, don't merge.

The story behind the numbers

If the figures above are the ledger, this is what they mean — ten short stories, then five paragraphs. (Full text: 00.0-stepping-back-from-CMA-operations/project-stories.md)

  1. The tower of paper. Four scans, testing the agenda-retrieval software on 1,200 places, 18,058,893 pages across 1.4 million PDFs and 2.18 terabytes — the quiet, unglamorous bulk of American local democracy, the meetings almost nobody attended, preserved anyway.
  2. The past is where the rot lives. Of the whole haul's 20,474 unreadable files, 19,418 sit in the single oldest corpus — the 2022 CivicPlus forensic set. The newer runs are nearly clean. Decay isn't hypothetical; it has an address, and it's the archive you didn't get to first.
  3. You can read a city's soul in its mean page count. Legistar averages 6.6 pages a document; CivicPlus averages 30.2. One vendor's world is terse indexes and action summaries; the other's is fat agenda packets. Same democracy, two dialects of paperwork.
  4. The board that never existed. In Texas, 6,613 documents were filed under a body called “General” — because the software defaulted an uncategorized meeting to that word, and 106 towns never changed it. Four separate councils dissolved into one folder. A field can be full and still say nothing.
  5. The meeting no one wrote down. 11,553 meetings produced zero files and were recorded nowhere but a scrolling log — not a failure, not a match, just gone. The gap you can't count is the one that quietly becomes a statistic.
  6. Index is not inventory. The Legistar API will happily list a document it then refuses to hand you. The whole method turns on one rule learned the hard way: “we didn't look” must never be filed as “it isn't there.”
  7. Honesty, counted. 418,871 of 418,871 files traced back to a source URL and event id; the 1,918 that couldn't be retrieved were each given a name and a reason — a 404, a bare filename, a withdrawn doc — instead of being lumped together as loss.
  8. A map that didn't exist. Enumerating vendor subdomains from the Wayback index turned a working list of 2,414 into 12,362 tenants across ten platforms — a jurisdiction-to-vendor atlas of who runs American civic meetings that, as far as anyone knows, had never been drawn.
  9. The address in the permit. Buried in the fat packets ride names and home addresses from permit applications — the ethics problem of republishing the public record in bulk, where “technically already public” meets “now trivially searchable.”
  10. The clerk who knows every voice. The endgame isn't a scraper; it's Archive Corps — a retired city clerk who recognizes every council member by sound, a librarian who builds the subject vocabulary, volunteers doing the enrichment machines can't.

Somewhere in the United States tonight, a city council is meeting in a half-empty room, and a camera nobody is watching is recording it. Multiply that by tens of thousands of towns and forty years, and you have one of the largest, least-studied records of a democracy anywhere — and one of the most fragile, because it lives on volunteer laptops, dying platforms, and vendor servers that forget. This project is the attempt to catch it before it goes. The tally so far is deceptively dull and secretly enormous: eighteen million pages, one-point-four million documents, two-plus terabytes of agendas and minutes — the connective tissue of self-government, gathered so it can't be quietly deleted.

The most haunting number isn't the biggest one. It's that almost every unreadable file in the entire collection — nineteen thousand of them — belongs to the oldest corpus, the 2022 forensic set. The newer scrapes come back clean; the old one comes back with holes. That's the whole argument for the work in a single statistic: digital records don't fail all at once, they fail at the edges first, and every year you wait, the edges move inward. The forensic bundle isn't just data; it's a memento mori for the born-digital record.

Look closer and the corpus starts to gossip. Legistar towns publish lean, six-page index documents; CivicPlus towns publish thirty-page packets stuffed with backup material. The page-count histogram is really an ethnography — you can see which governments treat the agenda as a pointer and which treat it as the whole file, which is why the same tooling had to learn ten different vendor dialects, and why merely finding the towns meant drawing a 12,362-tenant map that didn't previously exist.

But the real story is epistemological, and it's the one worth telling twice. Again and again the project ran into the same trap wearing different masks: a document the API lists but won't deliver; a meeting that generated no files and so was recorded nowhere; six thousand records filed under a board named “General” that turned out not to be a board at all. Each time, the lesson was the same — a populated field is not a usable value, and silence is not absence. The discipline that came out of it is almost moral: count the thing itself, give every gap its own reason code, and reconcile in both directions until zero files are of unknown origin. That's why “418,871 of 418,871” is the proudest number here. It's not that nothing was lost — nearly two thousand documents were — it's that nothing was lost anonymously.

And it all points somewhere human. The endgame is this contribution to Democracy's Library and a three-tense architecture that reads more like a clock than a schema — collector.db for what the fleet is doing now, archive.db for what's already preserved, civic.db for every government that should be represented but isn't yet. The last database is a to-do list for a democracy's memory. Filling it won't be done by scrapers alone; it'll be done by Archive Corps — the retired clerk who knows every voice, the student learning archival standards, the neighbor who can name the people in a 2009 zoning hearing. The stories in these numbers are about decay and discipline, but the one they're all building toward is about stewardship: a community learning to keep its own record, at a scale no one has tried before.