Developer Rants: Current SOTA Models Are Practically Worse Than Last Gen
zeeg · x · 2026-07-28
Developer Keith Whor expressed strong skepticism regarding the actual performance of current state-of-the-art (SOTA) AI models.
- Capability Regression: He argues that recent top-tier models (cited as Opus 5, GPT-5.6) are not materially better than the previous generation, and might actually be worse.
- Harder to Control: These models frequently go out of scope during tasks, requiring significant planning and babysitting from developers to reign them back in.
- Increased Overhead: The need to constantly correct the model's deviations ultimately leads to much higher token consumption and maintenance effort.
More from Models
- ProteinGym-LLM ranks Claude Opus 5 highest on a 217-task protein variant benchmark — LeoTZ03 · 2026-07-28
- Kimi K2 Third-Party Inference Pricing Matches Official API as Location Becomes Key — kevinsxu · 2026-07-28
- Arena’s new factuality ranking puts Claude Opus 5 Max at #1 — arena · 2026-07-28
- Kimi K3 used 51.2 million sandboxes across 1.5 million images — tarantulae · 2026-07-28
- A parody chart claims GPT wins because its version number is larger — Tystros · 2026-07-28
- Claude Opus 5 looks strongest in model-welfare tests, but may just be best at taking them — TheZvi · 2026-07-28