Gemini 4 Argon posts lowest hallucination rate (15%) on AA-Omniscience benchmark

import_jmr · x · 2026-10-02

On Artificial Analysis' AA-Omniscience benchmark, Gemini 4 Argon shows the lowest hallucination rate among leading models: 15%, vs Grok 4.7 at 29%, GPT-6 Astra at 45%, and Opus 5.5 at 59%. The author argues the metric matters beyond error counts: lower hallucination suggests better metacognition — the model knows where its knowledge ends, which is what autonomy rests on.

Related event: Gemini 4 Argon Leads AA-Omniscience with 15% Hallucination Rate(2 posts)→

Original post →

More from Models

Models channel →