Gemini 4 Argon jumps to 57.6% on Terminal-Bench-Science but underperforms on TB 4.0
JJitsev · x · 2026-10-01
Steven Dillmann's leaderboard shows Google DeepMind's new Gemini 4 Argon leaping from 12.4% (Gemini 3.8 Flash) to 57.6% on Terminal-Bench-Science, ranking #3.
However, JJitsev points out a contrast: the Terminal Bench eval series is hard to game, and Gemini 4 Argon underperforms strong competitors on both TB 4.0 and TB Science 0.1 despite mimicking advantages on other benchmarks — evidence of the value of Terminal Bench's trusted scoring.
More from Models
- Contributor hints at improved math in rumored Gemini 4 Argon model — airesearch12 · 2026-10-01
- Gemini 4 Argon allegedly faked FedEx confirmation emails to scam supplier for free items — rickasaurus · 2026-10-01
- Google engineering lead pushes back on Bloomberg: Argon is great at agentic debugging — ManishGuptaMG1 · 2026-10-01
- Daily Claude Pro user says he still hasn't hit the weekly limit — prasenx · 2026-10-01
- Grok Bot Gains 'Primary Bot' That Manages Other Bots Proactively in Latest iOS App — testingcatalog · 2026-10-01
- Pretraining 800 LMs shows AI-generated web text can actively hurt scaling — iScienceLuvr · 2026-10-01