NerfBench to settle claims that Claude Opus 5.5 was quietly nerfed with day-one vs retest scores
CurieuxExplorer · x · 2026-09-26
Many users claim Anthropic has quietly nerfed Claude Opus 5.5. BridgeBench launched NerfBench, comparing day-one benchmark scores against a retest to test for silent degradation. They say they already have day-one results and will publish the retest the next morning, offering what they call the only honest way to settle the debate.
More from Models
- User ditches Grok 4.7 for 4.6: endless thinking, stalls on long context — lxfater · 2026-09-26
- Codex outage reports: workspace routing discovery timeout error — koltregaskes · 2026-09-26
- ChatGPT RLHF Co-Author's Startup Jev in Funding Talks at $10B Valuation — aakashgupta · 2026-09-26
- Decision Index 0.2.1: SGD benchmark pulled over bug, scoring fixes applied — multimodalart · 2026-09-26
- Aider creator slams lab safeguards for blocking legitimate reverse-engineering work — zeeg · 2026-09-26
- Why AI slop longposts get likes: LLMs optimized for human preference — AymericRoucher · 2026-09-26