AA-Omniscience chart ranks 28 models on accuracy and hallucination
teortaxesTex · x · 2026-07-21
A new AA-Omniscience comparison chart ranks 28 of 446 models on two axes: accuracy and hallucination rate.
Key takeaways
- In accuracy, Claude 4 Opus, GPT-5.1, Sol, and GPT-5.5 lead the chart, with the top model at 61%.
- In hallucination rate, the best models sit in the 14%–29% range, while several lower-performing models climb above 80%.
- The post’s takeaway is that the benchmark suggests a strong relationship between scale, model family, and “omniscience”-style reliability, but also exposes large variance within and across vendors.
The author frames the chart as evidence that some models are much better at knowing what they know, while others are notably more prone to confident errors.
Related event: AA-Omniscience Chart Compares Accuracy and Hallucination Rates of 28 Models(2 posts)→
More from Models
- Moonshot’s Kimi K3 reaches #5 on MathArena as the top open model — xeophon · 2026-07-22
- Google launches Gemini 3.5 Flash Cyber for CodeMender, with limited access for governments — GoogleAI · 2026-07-22
- Kimi K3 tops Gemini 3.6 Flash on four shared public benchmarks — ChrisGPT · 2026-07-22
- Google’s year-long pause in new base-model pretraining draws sharp criticism — teortaxesTex · 2026-07-22
- Current setup is 8,192 input tokens and 2,048 output tokens, with 8k/512 next — TheZachMueller · 2026-07-22
- Kimi K3 feels slower than K2.7, but stronger on long coding jobs and refactoring — Far-Presence2711 · 2026-07-22