WeirdML v2 Benchmark Released: Tasks Expanded to 19, Clear Cost-Performance Scaling

yacineMTB · x · 2026-08-07

WeirdML v2 benchmark has been officially released, featuring significant expansions and upgrades:

Additionally, the tweet notes that Muse Spark 1.2 (xhigh) scored 60.3% on the leaderboard, which is solid but still falls short of the frontier models.

Original post →

More from Models

Models channel →