If humans can't tell new models apart, does bench maxing become labs' best investment?
Tim_Apple_938 · reddit · 2026-10-05
A Reddit user proposes that baseline quality of new models is now so good humans can't reliably tell them apart without citing benchmarks — e.g., why is Astra better than Fable, or Opus 5.5 than 6.1 Sol?
The provocative corollary: if perception can't differentiate models, benchmark maxing may become the most rational spend for private labs, since scores are the only visible signal of improvement. The thread debates whether evals are replacing lived experience as the core competitive metric.
More from Models
- Claude's self-proving consciousness answer: it reenacted the 5 tests scientists use on babies and animals — VoidStateKate · 2026-10-05
- Dev asks Claude Opus 5.5 to visualize the inside of its own mind — ZeroStateReflex · 2026-10-05
- 'All the butterflies will stay dead forever': viral jab at AI art homogenization — moultano · 2026-10-05
- Users want personality and voice customization back as AI assistants feel bland — koltregaskes · 2026-10-05
- DiffusionGemma plays 2048 straight from pixels at <150ms p95 latency — spillai · 2026-10-05
- VLM spot-the-difference latency race: 150ms vs Cloudflare CLEF's 450ms+ — spillai · 2026-10-05