LiveNerf repo tracks Opus 5.5 daily benchmarks to detect if Anthropic nerfs the model
TheOnlyVibemaster · reddit · 2026-09-28
To test community claims that Opus 5.5 has been quietly nerfed, a developer is re-running independent benchmarks like GPQA and SWE-bench daily and graphing results in an open repo called LiveNerf. So far, no nerf detected. The methodology uses Claude to pick questions Opus 5.5 struggles with most, with Opus 5 as a control; a deviation over 7.5 points would signal potential capability loss. Baseline expected by day 10, real judgments deferred until day 20 — a data-driven answer to whether the 'nerfed' discourse is hysteria.
More from Models
- User finds Opus 5.5 usage on Claude a far better deal than Astra on Codex — RexDouglass · 2026-09-28
- Dev builds private "gamefeel" benchmark to test LLMs on game design intuition — teortaxesTex · 2026-09-28
- Blogger predicts all major US AGI labs will reach AGI in 2027, then ASI — Dr_Singularity · 2026-09-28
- Three underrated Gemini 3 Flash use cases: browser use, image gen, and grunt work — BuffaloConscious7919 · 2026-09-28
- Student trial ships no Pro quota and blocks Gemini 3.8 Flash despite UI claims — EmoLotional · 2026-09-28
- Course notes explore plugging calibrated System-1 models like Jev into probabilistic programming — xuanalogue · 2026-09-28