Hallucination benchmark: Gemini 4 Argon 15% vs GPT-6 Astra 51% and Opus 5.5 66%

brandon_galang · x · 2026-10-01

AA-Omniscience hallucination-rate results show a huge gap among flagships: Gemini 4 Argon at 15%, GPT-6 Astra at 51%, and Opus 5.5 at 66%. Brandon Galang argues that despite wariness about Gemini benchmaxxing, this makes Argon the go-to model for enterprise chat — saying "idk" beats fabricating answers for average users, and knowing when to shut up is the ultimate benchmark.

Original post →

More from Models

Models channel →