Bloomberg: Gemini 4 aces benchmarks but struggles at real coding tasks internally
firstadopter · x · 2026-10-01
Bloomberg reports that as Google prepares to launch Gemini 4, it faces internal skepticism about the flagship model's real-world performance.
- Sources say Gemini 4 scores well on industry benchmarks but underdelivers when employees actually put it to work
- The model reportedly struggles with certain coding tasks, per people with direct access
- Commenter firstadopter notes Google models have historically been bench-maxxed; real verdict awaits developer hands-on
A notable contrast between pre-launch hype and internal reality.
More from Models
- GPT-6.1 Sol benchmarks across all effort levels land, making GPT-6 Astra hard to justify — PawelHuryn · 2026-10-01
- Gemini element-naming race heats up: Neon and Argon taken, Krypton is next as Google rejoins the frontier — tkipf · 2026-10-01
- Rumor: Gemini 4 Argon will launch as Ultra subscriber exclusive at first — opmgyhx · 2026-10-01
- Researcher: Model's unprompted video-joke disclaimer is hard to explain without 'understanding' — technollama · 2026-10-01
- Talking math and physics with LLMs feels like working with an exam-acing savant — burny_tech · 2026-10-01
- One-line take: GPT-6.1 Sol is underrated, says AI commentator — mallow610 · 2026-10-01