Spanda: Open-Source Hallucination Detector Runs in 1.5ms on CPU, 90,000x Faster than Semantic Entropy
Otherwise_Nobody_721 · reddit · 2026-09-05
- Spanda (MIT, pip install spanda) is a zero-dependency, pure-Python hallucination detector positioning itself as a production alternative to Semantic Entropy (Farquhar et al., Nature 2024).
- The baseline's pain: Semantic Entropy needs DeBERTa clustering at 136 seconds per query plus GPU VRAM. Spanda instead samples K outputs and computes a normalized lexical consensus ratio: 1.5ms on CPU, zero extra API/VRAM cost—roughly 90,000x faster.
- Accuracy holds up: 0.889 AUROC on GSM8K for 7B–27B models, matching heavy NLI clustering.
- Key caveat discovered: "Confident Mode Collapse"—heavily aligned frontier models (120B+) can repeat the exact same hallucinated answer across all seeds even at temperature 0.7, at which point all self-consistency methods fail.
- Author invites discussion on real-time uncertainty scoring in production pipelines.
More from Research
- Homework for researchers: extending RoPE to tensor product representations — thomasahle · 2026-09-05
- Debate: Models Fuzzily Recall Concepts, Not Text — SAE Features vs Edit-Distance Memorization — voooooogel · 2026-09-05
- ICML Position Paper: Unlabeled Data Doesn't Mean No Human Supervision — serrjoa · 2026-09-05
- VLA-Corrector from ZJU & Alibaba DAMO lifts robot success rates while cutting policy calls — 机器之心 · 2026-09-05
- Bug Hunt Bench: 105 real bugs stress-test GPT-6, Claude, Grok, Gemini and more coding agents — PawelHuryn · 2026-09-05
- Many mathematicians value prestige over truth, discussion on AI proofs notes — avt_im · 2026-09-05