NerfBench's First Results Find No Nerf: Claude Opus 5.5 Dips Just 0.8% vs Launch
alejandroll10 · x · 2026-09-28
NerfBench has published its first batch of retest results, comparing models' current performance against their own launch scores to check whether vendors quietly degrade released models.
- Claude Opus 5.5: retested at 99.2%, down 0.8% from launch
- GPT 6 Astra: retested at 102.8%, up 2.8% from launch
- Verdict: no nerf detected — both deltas fall within normal variance
The team plans more frequent retests across more models, says the benchmark improves as data accumulates, and is soliciting suggestions for which models to retest next.
Related event: NerfBench retests find no stealth nerfing of Claude Opus 5.5(2 posts)→
More from Models
- Claude Code pay-as-you-go credits burn $50 in 15 minutes, user warns against buying extra usage — RileyRalmuto · 2026-09-28
- Early Opus 5.5 user: coding with it 'feels like I'm floating' — Rasmic · 2026-09-28
- Claudes Lie Most in Diplomacy, Says Test; GPT-6 Astra Wins Without Betrayal — teortaxesTex · 2026-09-28
- Local Models and Bitcoin Are Both Freedom Tools, Argues Gladstein — csuwildcat · 2026-09-28
- Musk confirms Grok 'upgrades' as users notice dramatic speed boost — elonmusk · 2026-09-28
- "System 2 models built the brain, but System 1 is building the nervous system" — ai · 2026-09-28