Live SlopCodeBench run puts Opus 5 ahead of Opus 4.8 and Sonnet 5
HamelHusain · x · 2026-07-25
Opus 5 is being benchmarked against Opus 4.8 and Sonnet 5 on SlopCodeBench
A live benchmark run is comparing three Claude models on SlopCodeBench, with the image showing an early progress dashboard.
- At the time of the screenshot, 8 of 51 checkpoints were complete.
- Spend so far was $12.31, with a projected total of about $102.
- The dashboard showed Opus 5 leading on the displayed problem, with 2/8 strict passes already complete.
- The poster also reported early counts on the first problem: Opus 4.8 took 7 turns, Opus 5 took 11 turns, and Sonnet 5 had reached 33 turns and was still going.
This is an in-progress eval, so the takeaway is preliminary rather than final.
More from Models
- Claude Opus 5 builds a Rocket League clone on just 27% of a Max plan — soumitrashukla9 · 2026-07-25
- Opus 5 adds numeric self-checks to the Boeing benchmark and outbuilds Fable — victormustar · 2026-07-25
- Repost accuses Claude benchmark chart of highlighting only GPT-5.6 Sol’s sole win — soumitrashukla9 · 2026-07-25
- Opus 5 chart shows competitive gains across coding, search and computer use — CtrlAltDwayne · 2026-07-25
- Hands-on: A New Cost-Effective and Fast Model for Agent Workflows — doodlestein · 2026-07-25
- Claude Opus 5 reportedly beats Fable 5 on a hard 3D coding test at 75% of the price — rohanpaul_ai · 2026-07-25