Three frontier models in a row see users rolling back to older versions
alexisgallagher · x · 2026-09-12
- @kunchenguid tracks user sentiment and reports three consecutive frontier model launches where many users rolled back: Opus 5 (widely negative — poor communication, jumping to conclusions, repeat mistakes, even a dedicated complaint site; users returned to 4.6/4.8) and GPT 6 Astra (users reverting to 5.6 due to high cost as a daily driver and spiky performance).
- @alexisgallagher argues models are genuinely improving, but slower than user demands are growing: as models generate more code, users need them to handle global architectural coherence in larger codebases and parallel tasks — hence the focus on chief-of-staff patterns and embedding architectural invariants into repos.
- Some speculate excessive RL may explain the perceived slowdown in quality gains.
More from Models
- Codex quality fixes shipped, reset rolling out to ChatGPT Work users — CtrlAltDwayne · 2026-09-12
- Codex usage reset reportedly rolling out Sep 12 at 07:00 UTC, execution unconfirmed — DevDminGod · 2026-09-12
- Qwen3.8-27B goes live on Cerebras with fast inference, scoring 34 on AAII — Alibaba_Qwen · 2026-09-12
- antirez uploads DeepSeek V4.1 Flash GGUF quantizations to Hugging Face — Queasy_Asparagus69 · 2026-09-12
- Qwen3.8-Max scores 19 on OpenRouter vs 26 on Alibaba — effort level bug suspected — PawelHuryn · 2026-09-12
- Pipecat v1.9 adds Meta's Muse Voice Transcribe, the lowest semantic-WER STT model tested — solyarisoftware · 2026-09-12