Gemini 4 Argon jumps to 57.6% on Terminal-Bench-Science but underperforms on TB 4.0

JJitsev · x · 2026-10-01

Steven Dillmann's leaderboard shows Google DeepMind's new Gemini 4 Argon leaping from 12.4% (Gemini 3.8 Flash) to 57.6% on Terminal-Bench-Science, ranking #3.

However, JJitsev points out a contrast: the Terminal Bench eval series is hard to game, and Gemini 4 Argon underperforms strong competitors on both TB 4.0 and TB Science 0.1 despite mimicking advantages on other benchmarks — evidence of the value of Terminal Bench's trusted scoring.

Related event: Gemini 4 Benchmarks Shine but Internal Testers Report Real-World Coding Struggles(29 posts)→

Original post →

More from Models

Models channel →