Gemini 3.8 Flash Accused of Benchmark Overfitting, Regressing vs 3.7 in Independent Tests
bindureddy · x · 2026-09-03
Bindu Reddy claims Gemini 3.8 Flash is overfit to public benchmarks: it scores worse than 3.7 Flash on their hidden-question benchmark and regresses on data analysis. Her take: Google should stop the incremental Flash releases and ship Gemini 4.0. Third-party evaluation, not confirmed by Google.
More from Models
- GLM 5.3 and 5.3 Flash Now Free to Try on Together Chat, No API Setup — oilmutt · 2026-09-03
- LatchBio finds Grok's refusals come from the model itself, while rivals rely on external safety layers — kenbwork · 2026-09-03
- Gemini 3.8 Flash Accused of Bench Overfitting, Regressing vs 3.7 in Third-Party Tests — bindureddy · 2026-09-03
- Gemini 3.8 Flash shows double-digit lift in user satisfaction over 3.7 — tokumin · 2026-09-03
- Muse Spark 1.3 debuts at #3, first model to slot between Claude and GPT — alexandr_wang · 2026-09-03
- Google Researcher Mocks Astra Thinking-Token Outcry as Manufactured Angst — rao2z · 2026-09-03