Civic AI Audit
Grading what chatbots tell Americans about their own government.
- 2025–
- Conceived and developed the framework, question bank, and method
- live
- 5,560
- 3,761
- 20
People have started asking chatbots the questions they used to ask a search engine, a librarian, or a poll worker. Who represents me? When do I vote? What did my senator do? Nobody was checking whether the answers were right.
The Civic AI Audit checks. It measures the accuracy of consumer AI chatbots on civic questions, grading every captured response against time-versioned ground truth.
Why time-versioned matters
Civic facts expire. A representative loses an election, a deadline passes, a district is redrawn — and an answer that was correct in March is wrong in November. Grading a model’s answer against today’s truth would be unfair; grading it against no truth at all would be useless. So every ground-truth record carries the window in which it was true, and each response is graded against the facts as they stood when the question was asked.
The framework
The underlying measurement framework is organized around a civic triad — the Information Commons, the Social and Economic Fabric, and Power and Representation — with a question bank spanning each pillar. It builds on and is positioned relative to earlier work by Proof News, Caucus AI, and Forum AI.
How it works
Four stages, with a human in the loop at both ends where judgment actually matters:
Author. Define the axes of the audit matrix — question templates, value sets, and targets. A single template fans out into thousands of concrete test cases.
Ground truth. Import correct answers from official sources. Where sources disagree, the conflict goes to a human review queue rather than being resolved silently.
Run. Compose a template into test cases, fan out a batch, and dispatch execution jobs to runners that query the models.
Grade. An LLM judge proposes a verdict against ground truth. A human confirms or overrides it. The machine does the volume; the person owns the call.
What it produces
Dashboards for the Senate and the House, plus a running set of key insights about where models are reliable and where they confidently aren’t.