Gemini 4 Argon hallucinates half as often as rival frontier models on AA-Omniscience
zacharynado · x · 2026-10-06
- Artificial Analysis' AA-Omniscience benchmark (week of Oct 5) shows Gemini 4 Argon makes things up only 15% of the time when it doesn't know an answer — about half the rate of the next best frontier model.
- Behind it: Qwen3.8 Max (29%), Grok 4.7 (29%), and Muse Spark 1.3 (32%). GPT-6 Astra has the highest accuracy of the six (61%) but still guesses on 45% of the questions it misses.
- Whether a wrong answer or no answer costs more depends on the task; routing requests per task (e.g. via a gateway) is suggested.
More from Models
- Daniel Han publishes summary of LLM benchmarks you can actually trust — danielhanchen · 2026-10-06
- Claim Verification Benchmarks Mostly Test Retrieval, Not Reasoning, Finds 24K-Trace Study — deliprao · 2026-10-06
- COLM26 study: LLMs ace claim verification benchmarks by taking shortcuts, not verifying — deliprao · 2026-10-06
- Opus 5.5 uses 26k tokens vs Astra's 12k yet costs 23% less per task at equal AA score — ChrisGPT · 2026-10-06
- GPT-6 Astra claimed to be first AI crossing world-class astrophysics threshold — johnseach · 2026-10-06
- $500/mo AI subscription is huge money in Jakarta: PPP pricing debate — sujingshen · 2026-10-06