Gemini 4 Argon (High) tops Text Arena at 1525 pts, reshapes Pareto frontier at $0.62/task
arena · x · 2026-10-01
- Agent Arena: Google DeepMind's Gemini 4 Argon (High) ranks #8 with a +7.92% net improvement score at just $0.62 per task, reshaping the cost-quality Pareto frontier — a 4.96-point jump over Gemini 3.8 Flash (High) at #19.
- Text Arena: Per the quoted tweet, it takes #1 at 1525 pts, 20 points above #2 Claude Opus 4.6 (High), with a blended $8/MToken making it the most cost-efficient model; it leads Coding, Hard Prompts, Instruction Following, Longer Query, Creative Writing, plus every occupational domain including English, Chinese and Russian.
- Key signals: #1 Steerability (+15.88%), #2 Confirmed Success (+14.15%), #4 Praise vs Complaint (+27.74%); #3 in Chat (+11.58%).
- Scores are preliminary, based on 3k real-world agentic sessions so far. WebDev Code Arena: #8 (1679 pts).
Related event: Gemini 4 Argon Ranks 8th on Agent Arena(2 posts)→
More from Models
- One-line take: GPT-6.1 Sol is underrated, says AI commentator — mallow610 · 2026-10-01
- "Opus 5.5 is the cheapest model" — human hours saved beat token pricing — CamBrazy3 · 2026-10-01
- CUHK study maps when recurrence helps in looped language models, proposes history-state injection — CUHK-CSE · 2026-10-01
- Local 27B Face-off: Dirk-Qwen3.8 Beats Swift-1.5 on a 200-Question Personal Eval — norenEnmotalen · 2026-10-01
- Influencers hype Gemini 4 Argon: GOATED or hopelessly benchmaxed? — thatroblennon · 2026-10-01
- Reddit Pushback: OpenAI Users Subsidize Failed Experiments Like Atlas and Sora Via Price Hikes — dagerika · 2026-10-01