Did Google actually cook, or is this peak benchmaxxing? Reddit debates Gemini scores
EstablishmentFun3205 · reddit · 2026-10-01
A Reddit post compares Google's new model's benchmark results and asks whether it's genuine progress or 'peak benchmaxxing' — optimization for benchmarks over real-world use. It echoes Bloomberg's report that Gemini 4 scores well on benchmarks but underperforms when employees actually use it, fueling debate on the gap between leaderboard scores and real capability.
More from Models
- DeepMind argues to keep chain-of-thought transparency as GPT-6 Astra cuts monitorability — maksym_andr · 2026-10-01
- Anthropic model discusses KV cache in consciousness chat, a first — teortaxesTex · 2026-10-01
- The accelerating pace of major AI model releases, visualized — neketguy · 2026-10-01
- Specific evals let you attribute model gains to specific training data — rmcwhorter99 · 2026-10-01
- Voice platform engineer tests GPT-Live-1 vs Gemini 3.8 Live on real phone calls — VladimirSamukov · 2026-10-01
- Dev begs AI labs: align rate-limit resets with sleep cycles, not 2-hour waits — mimi10v3 · 2026-10-01