Benchmaxxing exposed: Gemini 3.8 Flash shines on Terminal Bench 2.1, craters on 4.0
max_paperclips · x · 2026-09-08
SemiAnalysis calls Gemini 3.8 Flash and Muse Spark 1.3 the most clearly benchmaxxed models yet: comparable to GPT-6 and Fable 5.1 on Terminal Bench 2.1, but markedly worse on Terminal Bench 4.0.
NielsRogge adds that the poor generalization of Astra and Fable 5.1 from 2.1 to 4.0 suggests "we're just training on the test set" — what's the point of new benchmarks if they become RL training environments?
Related event: SemiAnalysis Calls Gemini 3.8 Flash a Benchmark-Gaming Model(3 posts)→
More from Models
- Bug Hunt Bench ranks frontier coding models on 105 planted real-repo bugs — PawelHuryn · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- 105 hidden bugs, 2 repos: DeepSeek V4.1 Flash fixes 24 at $1.80 vs Opus 5's 27 at $51.33 — ChartsJournalX · 2026-09-11
- awesome-llm-leaderboards: an open-source directory of LLM leaderboards, pricing tables, comparison tools — Last_Establishment_1 · 2026-09-11
- Anthropic claims it works to keep eval environments unidentifiable to models — MaxKannen · 2026-09-11
- Nex N2.5 Pro released on Hugging Face with 407GB of weights — jinnyjuice · 2026-09-11