11 Frontier Models Take on Onchain Challenges: Best Scores Just 53%
davidfromkansas · x · 2026-09-02
Developer karinadoteth turned the Wintermute Alpha Challenge 2026 into a checker-verified benchmark testing 11 frontier and OSS models on blockchain coding and analysis — 8 public challenges, 750 points, with Foundry tests on historical chain forks and SHA-256-sealed answer keys.
Key finding: even with 2 hours of clock time and 200 tool calls, the best model (Opus 5) completed only 53% of the challenge; Grok and Fable also topped out at 53%. Models handled protocol coding via strong reasoning (Claude) or fast iteration (OSS), but all failed at onchain navigation and analysis — a capability cliff the author believes requires better training data or fine-tuning.
More from Models
- WSJ: Gemini 3.8 Flash drops tomorrow, preferred over Opus in internal coding tests — kimmonismus · 2026-09-02
- Fable 5.1 solves reading comprehension perfectly with zero reasoning — Sauers_ · 2026-09-02
- Fable 5.1 drains Claude 5-hour limit in under 30 minutes, users report — robleclerc · 2026-09-02
- Replit Announces Atlas: A Versatile Autoregressive Multimodal Model — gowthami_s · 2026-09-02
- Astra Model Touted as Impressive by Industry Observers — inductionheads · 2026-09-02
- Grok 4.6 and Fable 5.1 lead CursorBench pareto frontier — GavinSBaker · 2026-09-02