Gemini 4 Argon Tops Multiple Benchmarks, Matching GPT-6 Astra at Lower Cost
Google DeepMind has released Gemini 4 Argon, its first proprietary model in over 7 months to go beyond the Flash tier. Third-party benchmarks show it matching or even topping OpenAI's and Anthropic's flagship models across multiple benchmarks at significantly lower cost, signaling that the frontier model race has reached a dead heat.
Confirmed
- Artificial Analysis Intelligence Index: Gemini 4 Argon's high-reasoning tier scored 53, ranking 8th among 223 models (category median 26), tying GPT-6 Astra and Fable 5.1 and 1 point above GPT-6.1 Sol; per @ConsciousWarrior citing AA benchmark data, Gemini 4 scored the same as GPT-6 Astra while costing about 40% less.
- AutomationBench-AA: Ranked first at 77.5%, 6 percentage points ahead of Claude Sonnet 5.5 (max, 71.3%) (@ArtificialAnlys).
- Vals AI's Vals Index: Gemini 4 Argon took first place at 68.9%, the first time for a Gemini model (@burnytech citing Vals AI's announcement).
- Hallucination rate: Per @idg23 citing AA data, Gemini 4 Argon's hallucination rate is as low as 15%, versus Grok 4.7 at 29%, GPT-6 Astra at 45%, Opus 5.5 at 59%, and Fable 5.1 at 69%, reflecting a style of refusing to answer rather than fabricating.
- @cedricchee added: some users say it leads DeepSWE v1.1 on enterprise workflows, and with longer reasoning enabled, the output limit rises from 640K to 1 million tokens (this is user-reported and not yet officially confirmed).
Unconfirmed
- Coding ability in question: @Angaisb noted, as Argon topped the Vals Index, that its coding ability may still be weak; @cedricchee also mentioned coding performance varies by task.
Why it matters
- Gemini 4 Argon matches GPT-6 Astra at a lower price, a clear price-performance advantage (@vitaliychiley reports its cost efficiency falls between two generations of GPT models), which could reshape how enterprises and developers choose models. @balianone pointed out that while the gap with OpenAI and Anthropic has clearly narrowed, the overall leaderboard race is still undecided.
2026-10-01 ~ 2026-10-01 · 13 related posts
- Episode 1: Google DeepMind Unveils Frontier Model Gemini 4 Argon(2026-10-01, 52 posts)
- Episode 2: Gemini 4 Argon Tops Multiple Benchmarks, Matching GPT-6 Astra at Lower Cost(2026-10-01, 13 posts)
- Episode 3: Rumor: Gemini 4 Argon Can Output 1M Tokens in One Response(2026-10-01, 4 posts)
- Episode 4: Google Engineer Praises Gemini 4's Impressive Troubleshooting Skills(2026-10-01, 2 posts)
Primary sources
- Gemini 4 Argon tops AutomationBench-AA at 77.5%, 6 points ahead of Claude Sonnet 5.5 — ArtificialAnlys ·
- Gemini 4 Argon (High) benchmarks: 53 on AA Intelligence Index, ranks #8 of 223 — ArtificialAnlys ·
- Gemini 4 Argon hits 53 on AA Intelligence Index, matches GPT-6 Astra at 60% of the cost — cedric_chee ·
- Gemini 4 Argon Tops the Vals Index, but Coding May Still Be Its Weak Spot — Angaisb_ · 2026-10-01
- Gemini 4 Argon Takes #1 on Vals Index for the First Time at 68.9% — burny_tech · 2026-10-01
- [source] Gemini 4 Argon tops AutomationBench-AA at 77.5%, 6 points ahead of Claude Sonnet 5.5 — ArtificialAnlys · 2026-10-01
- Full benchmark results for Gemini 4 Argon with high reasoning released — ArtificialAnlys · 2026-10-01
- [source] Gemini 4 Argon (High) benchmarks: 53 on AA Intelligence Index, ranks #8 of 223 — ArtificialAnlys · 2026-10-01
- [source] Gemini 4 Argon hits 53 on AA Intelligence Index, matches GPT-6 Astra at 60% of the cost — cedric_chee · 2026-10-01
- Gemini 4 matches GPT 6 Astra on Artificial Analysis benchmark at 40% lower cost — Conscious_Warrior · 2026-10-01
- Gemini 4 Argon scores 53 on Artificial Analysis Intelligence Index, ties GPT-6 Astra — haider1 · 2026-10-01
- Unverified: Gemini 4 Argon reportedly scores 53 on AA Intelligence Index, 1M-token output — cedric_chee · 2026-10-01
- Gemini 4 Argon scores 53 on AA Intelligence Index, cost sits between GPT 6.1 Sol and GPT-6 Astra — vitaliychiley · 2026-10-01
- Gemini 4 Argon Posts 15% Hallucination Rate, Far Below GPT-6 Astra's 45% — i_dg23 · 2026-10-01
- Gemini 4 Argon closes gap with OpenAI and Anthropic on Artificial Analysis Index v4.3.2 — balianone · 2026-10-01
1 near-duplicate retellings: Conscious_Warrior