Gemini 3.8 Flash matches Fable 5.1 on benchmarks but is "garbage to use", dev complains
chandan1_ · x · 2026-09-09
A commenter argues Gemini 3.8 Flash now matches Fable 5.1 on benchmarks while being terrible to actually use, calling it benchmark maxxing — a fresh jab at the gap between leaderboard scores and real-world usability.
Related event: Gemini 3.8 Flash Divides Opinion: Strong Benchmarks, Poor Real-World Use(2 posts)→
More from Models
- GPT-6 Astra tops RSI-Exam at 0.5126, 18.4% above GPT-5.6 Sol — HuaxiuYaoML · 2026-09-09
- APEX-Agents 1.1 benchmark update: Claude Fable 5.1 tops leaderboard at 68.6% — amaarora · 2026-09-09
- Frontier labs must shrinkflate the $200 subscription to upsell you to API pricing — StewartalsopIII · 2026-09-09
- User math: peak pricing 2x but base rate better, 252M cache tokens cost just $0.75/day — teortaxesTex · 2026-09-09
- Astra noticeably worse than Sol in long threads, dev finds handoff workaround — jdjohnson · 2026-09-09
- VC communism is over: frontier models on rationing force hard model-choice thinking — StewartalsopIII · 2026-09-09