HalluWorld Benchmark, Accepted at NeurIPS, Shows Models Nail Perception but Fail Simulation
xennygrimmato_ · x · 2026-09-29
HalluWorld is a controlled benchmark and methodology placing LLM agents inside fully specified reference worlds—grid worlds, chess, and terminal tasks—to verifiably measure and induce hallucinations. Accepted at NeurIPS 2026 Evals & Datasets.
Key findings:
- Perception is near-solved: 7 of 12 models hit 0.0% perceptual error in grid worlds; 5 reach 0.0% on chess without FEN provided.
- Simulation is not: every model hallucinates on at least 16.7% of grid memory probes, and the best models still fail 18–32% of causal chess probes when a wrong FEN is in context.
- Abstention lags: no model gets below 24% hallucination on grid uncertainty probes; realistic terminal tasks show 9–27% uncertainty hallucinations while other categories approach zero.
The takeaway: hallucination is not one capability—perception, causal simulation, memory, and uncertainty need separate measurement.
More from coding & agent
- After 15 years and 90 Zendesk triggers, one founder replaced it all with an AI agent — clemnt · 2026-09-29
- IQuest-Q1 open-sourced: 320B MoE with 15B active, 524K context, day-0 vLLM support — vllm_project · 2026-09-29
- As agents go always-on, developers must design a 'night mode' for safety — sujingshen · 2026-09-29
- Manus Founder Red Calls Agents 'People' — But Where's the Line on Delegated Authority? — sujingshen · 2026-09-29
- Muse vs Grok Bot vs Cue: identity, permissions and payments separate AI agents — sujingshen · 2026-09-29
- 105 real bugs benchmarked: Sonnet 5.5 max scores 55.5, beating GPT-6 Astra at 45 — PawelHuryn · 2026-09-29