WeirdML v3 launches: agentic benchmark with 11 hand-made ML tasks

scaling01 · x · 2026-09-18

Developer htihle released WeirdML v3, a fully agentic ML benchmark featuring 11 complex hand-made tasks. Models must explore unfamiliar data, build their own ML and data-analysis pipelines, and produce results despite limited data, unspecified goals, and very little feedback.

The earlier v2 expanded to 19 tasks with API cost tracking; results show performance scaling with cost, and a highly varied Pareto frontier where 11 models from 6 companies each hold the best accuracy-per-cost in some range.

Original post →

More from Models

Models channel →