Gemini 4 aces benchmarks but struggles with real work, Google employees say
Neurogence · reddit · 2026-10-01
Bloomberg reports that despite strong benchmark results, Google employees with direct access are skeptical of Gemini 4 in practice — the model struggles with certain coding tasks, per anonymous insiders. The poster adds that by consumer release, Anthropic and OpenAI may have already shipped their next-generation models.
Related event: Gemini 4 Shines in Benchmarks but Struggles Internally: Bloomberg(4 posts)→
More from Models
- Gemini 4 Argon reportedly live as Google's model cadence accelerates, unconfirmed — Dr_Singularity · 2026-10-01
- webAI's 3.6B TwIL-LM3-Pro beats VibeThinker-3B by 35% on formal logic, runs locally in 2GiB — rohanpaul_ai · 2026-10-01
- Quick benchmark: Sol 6.1 inference is nearly 6x slower than Opus despite better token efficiency — RexDouglass · 2026-10-01
- Reddit Users Question AA Intelligence Index After Sonnet 5.5 Outranks Fable 5.1 — Ill_Distribution8517 · 2026-10-01
- Creator uses Opus 5.5 to co-direct a short vignette set in Tokyo — ebbyamir · 2026-10-01
- Researcher jokes: don't use 'eternally confused' chatbots for nuclear crisis decisions — examachine · 2026-10-01