Bloomberg: Gemini 4 shines on benchmarks but struggles with coding in real use, insiders say
kimmonismus · x · 2026-10-01
Bloomberg reports that insiders with direct access say Gemini 4 performs well on widely used benchmarks but underdelivers in actual employee use, struggling with certain coding tasks. KOL kimmonismus amplified the report with 'we all hope these rumors aren't true', fueling debate over the gap between Google's leaderboard scores and real-world performance.
More from Models
- DeepMind argues to keep chain-of-thought transparency as GPT-6 Astra cuts monitorability — maksym_andr · 2026-10-01
- Anthropic model discusses KV cache in consciousness chat, a first — teortaxesTex · 2026-10-01
- The accelerating pace of major AI model releases, visualized — neketguy · 2026-10-01
- Specific evals let you attribute model gains to specific training data — rmcwhorter99 · 2026-10-01
- Voice platform engineer tests GPT-Live-1 vs Gemini 3.8 Live on real phone calls — VladimirSamukov · 2026-10-01
- Dev begs AI labs: align rate-limit resets with sleep cycles, not 2-hour waits — mimi10v3 · 2026-10-01