Gemini 3.8 Flash Accused of Bench Overfitting, Regressing vs 3.7 in Third-Party Tests
bindureddy · x · 2026-09-03
Developer bindureddy reports that Gemini 3.8 Flash appears overfit to public benchmarks: on his private hidden-question set it scores worse than 3.7 Flash, with a regression in data analysis. He still calls 3.7 Flash an excellent model and urges Google to stop shipping Flash increments and release Gemini 4.0 instead.
Related event: Developer claims Gemini 3.8 Flash is overfit to public benchmarks(2 posts)→
More from Models
- ML researcher: LLM docs cram 3-4 idioms per sentence, ruining readability — ZeeshanZiaML · 2026-09-03
- Brockman pitches proactive AI agents; critics mock 'buy concert tickets' demos — max_paperclips · 2026-09-03
- Gemini 3.8 held its Pareto frontier spot for just 3.5 hours before Muse Spark 1.3 undercut it — giffmana · 2026-09-03
- OpenAI historically favors Thursdays — will rumored "Astra" launch tomorrow? — D3VAUX · 2026-09-03
- Muse Spark 1.3 ships with an underrated result, one-line curl install for Muse Code — alexandr_wang · 2026-09-03
- Will AI labs start shipping nightly model checkpoints? — intellectronica · 2026-09-03