WeirdML v3 launches: agentic benchmark with 11 hand-made ML tasks
scaling01 · x · 2026-09-18
Developer htihle released WeirdML v3, a fully agentic ML benchmark featuring 11 complex hand-made tasks. Models must explore unfamiliar data, build their own ML and data-analysis pipelines, and produce results despite limited data, unspecified goals, and very little feedback.
The earlier v2 expanded to 19 tasks with API cost tracking; results show performance scaling with cost, and a highly varied Pareto frontier where 11 models from 6 companies each hold the best accuracy-per-cost in some range.
More from Models
- PrismML launches ternary Bonsai 2 27B: 5.9GB, 9x smaller, retains 98.2% of full-precision performance — airesearch12 · 2026-09-18
- Grok Voice Transcribe 2.0 hits 92.9% accuracy on phone calls at just $0.10/hour — XFreeze · 2026-09-18
- Meta's 'Muse Spark' spotted on Hugging Face, release appears imminent — eliebakouch · 2026-09-18
- Silent Model Revisions Break Everything — Versioning and Model Cards Are a Mess — xeophon · 2026-09-18
- Suspected Gemini 4 Pro generates animated peacock-on-a-bike SVG in Arena — CodeByPoonam · 2026-09-18
- Suspected Gemini 4 Pro builds a Three.js rocket from a single-line prompt — CodeByPoonam · 2026-09-18