Artificial Analysis launches AA-Omniscience: all but 3 models hallucinate more than they answer right
geoffwolfe · x · 2026-09-23
Artificial Analysis announced AA-Omniscience, a new benchmark for knowledge and hallucination spanning 6,000 questions across 42 topics in 6 domains. Scoring: correct +1, wrong -1, abstention 0 — making uncertainty management part of the benchmark rather than a loophole, fixing the incentive flaw of prior evals where wrong answers carried no penalty.
Headline finding: all but three models are more likely to hallucinate than give a correct answer on difficult questions. The benchmark complements the Artificial Analysis Intelligence Index. The post also notes embedded knowledge matters for tool use (a model should know MCP means Model Context Protocol before deciding to search), and ReasonCore says it uses AA-Omniscience-style tasks in training inventory for factual reliability, calibrated abstention, and domain knowledge.
More from Models
- Show today's LLMs to experts 10 years ago and they'd call it AGI — JacksonKernion · 2026-09-23
- Tired of running out of credits, this Reddit user says cheap model swarms work surprisingly well — AnotherWallace · 2026-09-23
- User complains Anthropic's new model auto-routes to fable, stays locked on cyber — BLUECOW009 · 2026-09-23
- Unconfirmed: Qwen4-35B-A3B reportedly being tested, not yet announced — AIFlow_ML · 2026-09-23
- Unverified Claim: Opus 5.5 Was Distilled From a Bigger Internal 'Teacher' Model — ivan_bezdomny · 2026-09-23
- Is Slop Dead? Hamel Husain tests writing with Opus 5.5 — HamelHusain · 2026-09-23