If humans can't tell new models apart, does bench maxing become labs' best investment?

Tim_Apple_938 · reddit · 2026-10-05

A Reddit user proposes that baseline quality of new models is now so good humans can't reliably tell them apart without citing benchmarks — e.g., why is Astra better than Fable, or Opus 5.5 than 6.1 Sol?

The provocative corollary: if perception can't differentiate models, benchmark maxing may become the most rational spend for private labs, since scores are the only visible signal of improvement. The thread debates whether evals are replacing lived experience as the core competitive metric.

Original post →

More from Models

Models channel →