Gemini 4 Argon tops deepswe at 77.9% and automationbench, still locked to trusted testers
weswinder · x · 2026-10-10
Reports say Google's Gemini 4 Argon is #1 on deepswe and automationbench, scoring 77.9% and 51.3% respectively.
Yet it remains locked behind Google's "trusted testers" program. Meanwhile the author notes nobody he knows is leaving Claude Code or Codex for Antigravity: "benchmarks literally don't matter. Just ship it."
Takeaway: a promising model stuck in gated testing, and a coding-agent market where real adoption clearly diverges from leaderboard wins.
More from coding & agent
- Indie dev builds a GT5-grade browser racing sim in ThreeJS, tuned to IMSA GT3 rulebook — AIandDesign · 2026-10-10
- Agent-built system hits 2,242 tok/s on AMD MI300As, 2.33× faster than SGLang in 105 hours — bariskasikci · 2026-10-10
- Seroter's Daily Reads: Agent-Friendly APIs, LLM Judges, and Ambient Quality Agents — rseroter · 2026-10-10
- Steve Yegge: Claude just reads the binary when it wants to know how closed-source Rex works — Steve_Yegge · 2026-10-10
- YC hosts an agent-first hackathon in San Francisco, sponsored by Supabase — ycombinator · 2026-10-10
- Running a fleet of long-horizon LLM agents: 9 keep-alive patterns and 5 unsolved problems — milkygirl21 · 2026-10-10