Detect LLM hallucinations in 1.3µs on CPU — but 120B models hallucinate with unanimous false certainty

More_Slide5739 · reddit · 2026-10-02

Local-model users often detect hallucinations with Oxford's Semantic Entropy (Nature paper), which needs an NLI cross-encoder doing pairwise comparisons — costly on consumer GPUs. The author built Spanda (Rsc), a zero-dependency Python metric: normalize strings from 5 samples at T=0.7 and compute exact-match Shannon entropy on pure CPU in 1.3 microseconds, zero GPU cost. Benchmarked on GSM8K/TriviaQA across 1.5B–120B models:

Practical takeaway: for 7B–27B models on structured tasks (math, code, JSON, SQL, discrete QA), 5-path exact-match entropy gives 0.89 AUROC with no neural guardrails. Code, logs and full writeup are open source on GitHub (Spnda).

Original post →

More from Infra

Infra channel →