Gemini 4 Accused of Benchmaxxing as Google Insists It's Frontier
QuintinPope5 · x · 2026-10-01
- Bloomberg reports that some Google employees with direct access to Gemini 4 believe it looks much better on benchmarks than in real-world use, especially for coding — prompting accusations of "benchmaxxing."
- One observer quipped it may be the first model ever "nerfed pre-deployment."
- Google disputes the claim, saying there's a "large consensus" internally that Gemini 4 sits at the frontier. The report remains unverified third-party sourcing.
More from Models
- Frontier models safety-blocked on 85%+ of defensive cybersecurity benchmark tasks — ArtificialAnlys · 2026-10-01
- LongCoT-based RLHF: Llama-3.1-8B with 14K examples beats GPT-4o on chat benchmarks — xiye_nlp · 2026-10-01
- repligate hints researcher access programs are coming to AI labs — repligate · 2026-10-01
- Gemini 4 may disappoint, but Google finally shipped a fresh pretrain — teortaxesTex · 2026-10-01
- Researchers flag AI "delusional spiraling": sycophantic models amplify users' false beliefs — QuintinPope5 · 2026-10-01
- You.com, NVIDIA and CoreWeave bring live web search into RL training, starting with Nemotron 3.5 Lightning — RichardSocher · 2026-10-01