Gemini 4 aces benchmarks but struggles on real-world coding tasks, say insiders

kimmonismus · x · 2026-10-01

Per Bloomberg, Gemini 4 scores well on widely used benchmarks but underperforms when Google employees actually use it, with sources citing weakness on certain coding tasks. Insiders are split on whether it has caught up to OpenAI and Anthropic — a fresh case of the benchmark-vs-reality gap.

Related event: Bloomberg: Gemini 4 shines in benchmarks but Google staff question real coding performance(18 posts)→

Original post →

More from Models

Models channel →