Hallucination benchmark: Gemini 4 Argon 15% vs GPT-6 Astra 51% and Opus 5.5 66%
brandon_galang · x · 2026-10-01
AA-Omniscience hallucination-rate results show a huge gap among flagships: Gemini 4 Argon at 15%, GPT-6 Astra at 51%, and Opus 5.5 at 66%. Brandon Galang argues that despite wariness about Gemini benchmaxxing, this makes Argon the go-to model for enterprise chat — saying "idk" beats fabricating answers for average users, and knowing when to shut up is the ultimate benchmark.
More from Models
- NormViz benchmark: best model Gemini 3 Flash scores just 25.3% on visual cultural norms across 16 countries — StellaLisy · 2026-10-01
- Leak: OpenAI quietly added MCP events support at DevDay, enabling email subscriptions without polling — banteg · 2026-10-01
- LangChain launches LangSmith Fine-Tuning: turn agent traces into fine-tuned models via smithtune CLI — LangChain · 2026-10-01
- Reddit user: GPT-5.6 Sol nerfed so hard it needs 3 tries to swap a font, at 3x token price — Existing-Slide7395 · 2026-10-01
- "Pro" model that can't clearly beat GPT and Claude, locked behind a three-tier rollout — taiwbi · 2026-10-01
- Early take on Gemini 4 Pro: Astra-level 3D games, fewer hallucinations than Opus — bindureddy · 2026-10-01