AA-Omniscience chart ranks 28 models on accuracy and hallucination
teortaxesTex · x · 2026-07-21
A new AA-Omniscience comparison chart ranks 28 of 446 models on two axes: accuracy and hallucination rate.
Key takeaways
- In accuracy, Claude 4 Opus, GPT-5.1, Sol, and GPT-5.5 lead the chart, with the top model at 61%.
- In hallucination rate, the best models sit in the 14%–29% range, while several lower-performing models climb above 80%.
- The post’s takeaway is that the benchmark suggests a strong relationship between scale, model family, and “omniscience”-style reliability, but also exposes large variance within and across vendors.
The author frames the chart as evidence that some models are much better at knowing what they know, while others are notably more prone to confident errors.
Related event: AA-Omniscience Chart Compares Accuracy and Hallucination Rates of 28 Models(2 posts)→
More from Models
- RoMa v2 image matching model unveiled in the usual black poster — ducha_aiki · 2026-09-11
- OpenAI rated Astra 'Critical' for cyber capabilities — and admits it's harder to monitor — theguywhobuilds · 2026-09-11
- TestingCatalog's Daily AI Brief adds email editions, dishing Meta Muse and GPT-Live-1 rumors — testingcatalog · 2026-09-11
- ChatGPT monthly active users top 1.06 billion in August, fourth straight record month — FinanceYF5 · 2026-09-11
- PuzzleMask: Plain-Prose Attack Bypasses All 4 Tested LLM Gatekeepers at 100% — TechNadu · 2026-09-11
- OpenAI Codex may issue another usage reset this weekend, says Codex lead resets happen — umesh_ai · 2026-09-11