Reddit user: frontier models now feel indistinguishable in everyday use
rainyaltaccount · reddit · 2026-09-24
A Reddit user says he can barely tell frontier models apart anymore. Despite claims that Opus 5.5 is much better — and he does use 5.5 high — he sees little difference versus other top models in normal use.
He suspects his tasks may not be hard enough, and notes the clearest differences now come from the surrounding products, like how Codex vs Claude Code format responses. He asks whether most model gains now only show up on extremely hard benchmarks rather than everyday usage.
More from Models
- A 10-year trend holds: small fine-tuned models on selective data still beat bigger general models — xeophon · 2026-09-24
- GPT-6 Luna shows vision regression vs GPT-5.6: extraction drops 81.79% to 66.67% — ducha_aiki · 2026-09-24
- Luna 6 private coding evals don't look great — Maasu · 2026-09-24
- Claude Opus 5.5 tops Artificial Analysis at 58; four new models add 11 Pareto frontier points — ArtificialAnlys · 2026-09-24
- METR says it used an undisclosed 'additional source' to understand Anthropic's AI R&D, buried in the Opus 5.5 system card — coherence · 2026-09-24
- Hands-On: 6 Sol at Max Tier Underperforms Astra Light with Frequent Regressions — pwlot · 2026-09-24