Live SlopCodeBench run puts Opus 5 ahead of Opus 4.8 and Sonnet 5
HamelHusain · x · 2026-07-25
Opus 5 is being benchmarked against Opus 4.8 and Sonnet 5 on SlopCodeBench
A live benchmark run is comparing three Claude models on SlopCodeBench, with the image showing an early progress dashboard.
- At the time of the screenshot, 8 of 51 checkpoints were complete.
- Spend so far was $12.31, with a projected total of about $102.
- The dashboard showed Opus 5 leading on the displayed problem, with 2/8 strict passes already complete.
- The poster also reported early counts on the first problem: Opus 4.8 took 7 turns, Opus 5 took 11 turns, and Sonnet 5 had reached 33 turns and was still going.
This is an in-progress eval, so the takeaway is preliminary rather than final.
More from Models
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- awesome-llm-leaderboards: an open-source directory of LLM leaderboards, pricing tables, comparison tools — Last_Establishment_1 · 2026-09-11
- Anthropic claims it works to keep eval environments unidentifiable to models — MaxKannen · 2026-09-11
- Nex N2.5 Pro released on Hugging Face with 407GB of weights — jinnyjuice · 2026-09-11
- RoMa v2 image matching model unveiled in the usual black poster — ducha_aiki · 2026-09-11
- OpenAI rated Astra 'Critical' for cyber capabilities — and admits it's harder to monitor — theguywhobuilds · 2026-09-11