WeirdML v2 Benchmark Released: Tasks Expanded to 19, Clear Cost-Performance Scaling
yacineMTB · x · 2026-08-07
WeirdML v2 benchmark has been officially released, featuring significant expansions and upgrades:
- Task Expansion: The number of evaluation tasks has increased from the original 6 to 19.
- API Cost Tracking: It now tracks API costs and other metadata, providing deeper insights into the differences across various models.
- Cost vs. Performance: The cost vs. performance chart reveals a clear positive scaling trend (higher-cost models generally yield better results).
- Diverse Pareto Frontier: The benchmark shows a highly varied Pareto frontier, with 11 models from 6 different companies achieving top accuracy for specific cost ranges (e.g., Grok 3).
Additionally, the tweet notes that Muse Spark 1.2 (xhigh) scored 60.3% on the leaderboard, which is solid but still falls short of the frontier models.
More from Models
- ByteDance Training 10T-Scale Reasoning Model with 2000-Person Seed Team — alexvoica · 2026-08-07
- Qwen 3.8-Max Set to Drop Next Week: Starting with 2.4T Parameters, 27B to Follow — Ok-Shower7286 · 2026-08-07
- SK Telecom Releases A.X K2: A 688B Parameter Open-Source MoE Model — huggingface · 2026-08-07
- NVIDIA Exec: Reasoning Model Alpamayo Solves Self-Driving's Long Tail Problem — ZGojcic · 2026-08-07
- DeepSeek-V4-Flash Broken on AMD MI325X? User Reports Tool Call Chaos — Brunofcsampaio · 2026-08-07
- MameLoshnLM: First Open-Source 8B LLM and Benchmark for Yiddish — Yiddish-NLP · 2026-08-07