11 Frontier Models Take on Onchain Challenges: Best Scores Just 53%

davidfromkansas · x · 2026-09-02

Developer karinadoteth turned the Wintermute Alpha Challenge 2026 into a checker-verified benchmark testing 11 frontier and OSS models on blockchain coding and analysis — 8 public challenges, 750 points, with Foundry tests on historical chain forks and SHA-256-sealed answer keys.

Key finding: even with 2 hours of clock time and 200 tool calls, the best model (Opus 5) completed only 53% of the challenge; Grok and Fable also topped out at 53%. Models handled protocol coding via strong reasoning (Claude) or fast iteration (OSS), but all failed at onchain navigation and analysis — a capability cliff the author believes requires better training data or fine-tuning.

Original post →

More from Models

Models channel →