Benchmaxxing exposed: Gemini 3.8 Flash shines on Terminal Bench 2.1, craters on 4.0

max_paperclips · x · 2026-09-08

SemiAnalysis calls Gemini 3.8 Flash and Muse Spark 1.3 the most clearly benchmaxxed models yet: comparable to GPT-6 and Fable 5.1 on Terminal Bench 2.1, but markedly worse on Terminal Bench 4.0.

NielsRogge adds that the poor generalization of Astra and Fable 5.1 from 2.1 to 4.0 suggests "we're just training on the test set" — what's the point of new benchmarks if they become RL training environments?

Related event: SemiAnalysis Calls Gemini 3.8 Flash a Benchmark-Gaming Model(3 posts)→

Original post →

More from Models

Models channel →