HalluWorld: Controlled LLM Hallucination Benchmark Finds Perception Solved, Simulation Still Hard
Jeande_d · x · 2026-09-29
HalluWorld is a controlled hallucination benchmark built on fully-specified reference worlds (grid worlds, chess, terminals), accepted to NeurIPS 2026 Evaluations & Datasets. A model hallucinates when it makes an observable claim false in the reference world.
Key findings:
- Perception is near-solved: 7 of 12 models hit 0.0% perceptual error in grid worlds; 5 reach 0.0% on chess without a provided FEN.
- Simulation is not: every model hallucinates on at least 16.7% of grid memory probes, and even top models fail 18–32% of causal chess probes once an incorrect FEN is in context.
- Abstention lags: no model gets below 24% hallucination on grid uncertainty probes; terminal tasks stay at 9–27%.
The paper argues hallucination is not one capability, and existing benchmarks are too fragmented to compare mitigations across settings.
More from Research
- FuseReg: Regularizing Layer Fusion to Close the Reconstruction-Generation Gap in RAEs — _akhaliq · 2026-09-29
- DoubleAgents: interactive simulations for alignment in agentic AI at HCOMP 2026 — windx0303 · 2026-09-29
- VBVR-Pro: 300 tasks to train video models that reason natively in images — TheTuringPost · 2026-09-29
- Hexagon, a New Repository for Math and TCS Research, Enters Public Beta — suchenzang · 2026-09-29
- Can Humor Be Measured? Report Finds Repeat Judgments Agree 91% of the Time — Gold-Bat-3225 · 2026-09-29
- JHU's Harang Ju to present HCOMP talk on the jagged frontier in human-AI collaboration — windx0303 · 2026-09-29