LiveNerf hits ~1,000 stars tracking daily whether Opus 5.5 gets 'nerfed' post-release
TheOnlyVibemaster · reddit · 2026-10-02
A Reddit developer built LiveNerf, an open-source project that evaluates Opus 5.5 daily against a fixed benchmark to test whether models measurably degrade after release, instead of relying on anecdotes.
- The repo is approaching 1,000 stars, hit the Hacker News front page, and has drawn heavy scrutiny of its methodology.
- The project is still establishing a baseline; once enough data accumulates, statistical comparisons of later performance become possible. A finding of no degradation would also count as a result.
- Community criticism covers benchmark contamination, providers possibly detecting benchmark traffic, benchmark selection, and statistical methodology; the author welcomes issues.
- The author argues that since AI companies don't disclose model changes over time, the community should build its own measurement tools. Results will be posted once evidence emerges.
More from Models
- Fable 5.5 hallucinates less and replicates small models like Jev in hours, claims Bindu Reddy — bindureddy · 2026-10-03
- Ten Days of Nonstop Releases: Gemini, GPT, Sonnet and More New Models — haider1 · 2026-10-03
- Opus 5.5 and Sonnet 5.5 Now Available in Google Antigravity — NBMVegeta · 2026-10-03
- llama.cpp PR Halves Indexer Score Memory for Qwen Flash, Cutting VRAM Use — jacek2023 · 2026-10-03
- Prediction: Two Chinese Model Drops Coming, Both Claiming to Beat Opus 5.5 — bindureddy · 2026-10-03
- Anthropic's Opus 5.5 and Sonnet 5.5 Have Officially Arrived — Last_Conclusion_8984 · 2026-10-03