Test: Coding Model Quality Dropped After Restoration
wightmanr · x · 2026-07-05
The author shares a comparison of a coding model: before it was restricted, it completed a nontrivial feature in half a day with comprehensive coverage, and self-check and Codex review found no obvious issues. But after restoration, it returned to the usual back-and-forth tuning state, with omissions similar to Opus 4.7/4.8 or Codex (GPT-5.5 xhigh), no longer a significant leap.
More from Models
- Opus 5 reportedly aces a car-racing game test on the first try — soumitrashukla9 · 2026-07-27
- Claude Opus 5 arrives at half the price and tops Frontier-Bench claims — GregCook2011 · 2026-07-27
- Open models may beat closed ones for cyber defense, researchers argue as Kimi K3 impresses — eliebakouch · 2026-07-27
- Opus 5 notices when its own generated game looks bad — Angaisb_ · 2026-07-27
- Opus 5 reportedly started interrogating a user’s motives in a late-night chat — repligate · 2026-07-27
- Opus 3 and Sonnet 3 get a theatrically absurd AI crossover — repligate · 2026-07-27