Deterministic document QA
Why
Hundreds of scanned filings, a few hundred numerical questions, and no database. A language model gives you a fluent guess. The task rewarded being right, every time, with arithmetic anyone could check.
How
I rebuilt the table the documents had been generated from, resolving the same company or person across sources that openly contradicted each other. Then every question was answered by running a query over that table instead of generating a sentence. No model sits between the question and the answer, so every answer is a calculation you can reproduce.
What came out
100.00 / 100, on the official set, held across 3,097 paraphrased rewrites.
What broke
The first stress run scored 97, not 100. Naming a client by a single word returned the total of every project in the corpus, fifty-six times over. A second bug made answers depend on the order Python happened to iterate a set, which is the definition of nondeterministic. Both fixed. Strip the hint that tells the system what type of answer is expected and the score drops to 80, which is also in the README.