Claude Opus 4.8 missed full marks by tiny details on three ProgramBench tasks
jyangballin · x · 2026-07-25
The reply adds more detail on Claude Opus 4.8’s ProgramBench performance.
- The model came extremely close to a perfect rebuild three times, including near-misses on zip-password-finder, wrapcheck, and ngrrram.
- The misses were tiny: a quirky error string, a debug log line, or even a single keyboard shortcut.
- The point is that the benchmark failures are now down to very small implementation details.
Related event: Claude Opus 4.8 Hits Record 16.5% on ProgramBench(3 posts)→
More from Models
- Vercel adds Claude Opus 5 to AI Gateway with fast mode for coding agents — EricBuess · 2026-07-25
- Early Claude Opus 5 feedback says it helps ship PRs faster in Claude Code — EricBuess · 2026-07-25
- Claude Opus 5 feels like Fable, but cheaper — cedric_chee · 2026-07-25
- Anthropic’s Opus 5 looks more like a major upgrade than a minor refresh — yi_ding · 2026-07-25
- Anthropic launches Claude Opus 5, with blind tests placing it above GPT-5.6 — lennysan · 2026-07-25
- Anthropic says Claude Opus 5 was intentionally left untrained on cyber tasks — rez0__ · 2026-07-25