Pre-registered study finds a universal floor for hallucination detection, but no universal detector
k01234n · reddit · 2026-08-04
A pre-registered hallucination detector paper finds a universal floor, but not a universal detector
The author claims to have tested 10 models across two runs and four families of internal signals (29 signals total), with pre-registered evaluation done before seeing the data.
- Run 1: a geometry-only detector passed the pre-registered bar at 18/20 deployments. Adding the model’s own confidence did not improve the result, falsifying the stricter claim that confidence should cover more cases.
- There was no universal best signal: different signals won across the 18 working cases, so the best choice depends on the model and task.
- Still, a fixed combo calibrated on nine models and tested on the tenth beat chance on 9/10 ANLI and 10/10 TriviaQA, suggesting a universal floor rather than a universal detector.
- Run 2: on six additional tasks, calibrating per model reached 10/10 above chance.
- A single fixed blind detector, however, only got 6/10; in the four misses, the signal appeared inverted, with AUROC as low as 0.17.
The conclusion: the method can often find above-chance hallucination signals, but the sign and calibration still need to be set per model, so it is not a shippable universal detector. The paper also argues the signal is not a 4-bit quantization artifact, since the effect held across nf4 → int8 → bf16 → fp32 on the models where it worked.
More from Research
- A visual explainer breaks down what embedding models are and where they are used — _jaydeepkarale · 2026-08-04
- A neurosymbolic model is pitched as a way to build ultra-reliable numerical solvers — GaryMarcus · 2026-08-04
- Claude Code Under the Hood: System Prompts Exceed 75%, Size Surges — ClaudeCodeLog · 2026-08-04
- Self-distillation from production traces could make models improve with use — ypatil125 · 2026-08-04
- A new reading list links refactoring economics, agent skills, and Google’s Agent Skills — rseroter · 2026-08-04
- A well-executed PhD full of negative results should still be defendable — prajdabre · 2026-08-04