Interchange: change here between the Markets and ML lines.
Can a RAG system tell when it is about to be wrong
Why
A retrieval system that answers a question about a real company filing sounds exactly as confident when it is right as when it is wrong. A method called semantic entropy was published as a fix for that on general knowledge questions. I wanted to know if it actually works once the questions are financial and the source documents are real 10-Ks and 10-Qs, not trivia, and whether it beats something far cheaper.
How
I built a small retrieval pipeline over 150 real questions on 84 company filings from a published open benchmark, running a small open model on a laptop CPU, and had it answer every question five times. A completely different model graded each answer right or wrong before any detection method was scored, so the grading could not be shaped to flatter the result. Then I measured whether the disagreement across those five answers separates the right ones from the wrong ones, and lined it up against four simpler, cheaper signals doing the same job.
What came out
Semantic entropy clears its own pre-registered bar, but only just: the confidence interval bottoms out at 0.512, a hair above a coin flip. It does not beat a far simpler signal, exact-match agreement across the five samples, which scores 0.635, statistically indistinguishable from it. The model reads the correct page in 89 of 150 filings, and 29 of the 46 wrong answers happened with that correct page already sitting in front of it: the dominant failure is misreading a retrieved table, not failing to find it.
What broke
The headline comparison failed: semantic entropy was supposed to beat a cheap lexical shortcut and did not, 0.631 against 0.635, a gap the confidence interval calls noise. Of six numeric predictions written down before any answer was generated, three came in wrong, including the expected gain from trusting only the fifty most-certain answers, which fell far short. The sharper finding: on 15 of the 46 wrong answers, the model gave the identical answer all five times. It was completely consistent while being completely wrong, and a method built on disagreement across samples cannot see that failure by construction, not by bad luck.