Google models ace benchmarks but feel mid in use — will Gemini 4 Argon differ?
VraserX · x · 2026-10-01
Commentator VraserX calls out Google's recurring pattern: models look absolutely insane on benchmarks but feel mid in real-world use. He says Gemini 4 Argon looks incredible on paper and hopes this is the generation where Google finally proves his skepticism wrong.
More from Models
- Teknium: Hermes now supports adding new languages via plugins, on top of 16 built-in — Teknium · 2026-10-01
- Google ships Gemini 4 as contributors tout science and reasoning work — shagunsodhani · 2026-10-01
- Study: Language models hide 'bad news' in reports by default — Symbiot10000 · 2026-10-01
- Gemini 3.8 rates sexual assault less wrong when the victim is a man, test finds — OmegaGogeta · 2026-10-01
- Opus 5.5 tends to cross-contaminate separate projects in the same directory — repligate · 2026-10-01
- AI planned obsolescence: vendors hold a benchmark-invisible downgrade knob — maier_ak · 2026-10-01