Gemini 4 aces benchmarks but struggles on real-world coding tasks, say insiders
kimmonismus · x · 2026-10-01
Per Bloomberg, Gemini 4 scores well on widely used benchmarks but underperforms when Google employees actually use it, with sources citing weakness on certain coding tasks. Insiders are split on whether it has caught up to OpenAI and Anthropic — a fresh case of the benchmark-vs-reality gap.
More from Models
- DeepMind argues to keep chain-of-thought transparency as GPT-6 Astra cuts monitorability — maksym_andr · 2026-10-01
- Anthropic model discusses KV cache in consciousness chat, a first — teortaxesTex · 2026-10-01
- The accelerating pace of major AI model releases, visualized — neketguy · 2026-10-01
- Specific evals let you attribute model gains to specific training data — rmcwhorter99 · 2026-10-01
- Voice platform engineer tests GPT-Live-1 vs Gemini 3.8 Live on real phone calls — VladimirSamukov · 2026-10-01
- Dev begs AI labs: align rate-limit resets with sleep cycles, not 2-hour waits — mimi10v3 · 2026-10-01