Gemini 4 Scores Badly on cua-bench, Fueling Benchmaxxing Concerns
burny_tech · x · 2026-10-01
cua-bench creator spiceylemonade pushed back on Gemini 4's strong official scores: on their computer-use benchmark, Gemini 4 scores badly ("perhaps antigravity needs some work?"). The benchmark requires playing Minecraft, SUPERHOT, and unrevealed held-out games to prevent hillclimbing, though it's speed-bottlenecked at the further end.
Context: a cited report claims Gemini 4 performs well on industry benchmarks but underperforms when Google employees actually use it for work, reigniting "benchmaxxing" criticism. The gap between official scores and real-world agent performance is becoming a community flashpoint.
More from Models
- Local 27B Face-off: Dirk-Qwen3.8 Beats Swift-1.5 on a 200-Question Personal Eval — norenEnmotalen · 2026-10-01
- Influencers hype Gemini 4 Argon: GOATED or hopelessly benchmaxed? — thatroblennon · 2026-10-01
- Reddit Pushback: OpenAI Users Subsidize Failed Experiments Like Atlas and Sora Via Price Hikes — dagerika · 2026-10-01
- GPT 6 Series Shows Frequent Lazy Work, While Sol 5.6 Spent Hours on QA — jdjohnson · 2026-10-01
- Gemini-4-argon Debuts #1 on Arena's Text Leaderboard but Only 8th on WebDev — DeArgonaut · 2026-10-01
- Claim that DeepMind will beat Opus 5.5 at half price gets publicly called out as bogus — zacharynado · 2026-10-01