Benchmark Heaven Launches: Aggregates 100 Benchmarks, Ranks Models by Real Task Cost
On September 18, developer airesearch12 launched and open-sourced Benchmark Heaven (beta), an AI model evaluation aggregation site. The project is free, ad-free, and unfunded, with a core idea of making model comparisons closer to real-world usage costs and behavior.
Confirmed
- The site covers 100 benchmarks and 800 models, offering one-stop comparison and an overall Composite Score
- It introduces the novel Benchmaxxing score to measure a model's tendency to game benchmarks
- It proposes comparing models by "cost per real task": factoring in input/output tokens, inference time, provider filtering, and more — not just price per million tokens
- The project is open source; the author made two releases — the launch announcement (m1/m4) and an intro emphasizing the real-cost comparison angle (m2/m3/m5) — with consistent content
Why it matters
- The author's core argument is that "price per million tokens is a lie": a cheaper model that takes twice as long to think may actually cost more to complete a task. This per-task consumption-based pricing offers a fresh lens on evaluating the value of reasoning models
- The Benchmaxxing score directly addresses industry concerns about benchmark gaming and overfitting, quantifying "benchmark-gaming tendency" into a comparable metric
- Its free, open-source, ad-free, and unfunded positioning provides an auditable alternative to commercial evaluation sites
Timeline
- 09-18: Benchmark Heaven beta launched and open-sourced, with the author repeatedly introducing its aggregation capabilities and real-cost pricing approach
2026-09-18 ~ 2026-09-18 · 5 related posts
Primary sources
- Benchmark Heaven Launches: 100 Benchmarks, 800 Models, Real Per-Task Costs — airesearch12 · 2026-09-18
- [source] Benchmark Heaven aggregates 100 benchmarks and 800 models, adds first Benchmaxxing score — airesearch12 · 2026-09-18
- Price per million tokens is a lie: new benchmark scores models by cost per actual task — airesearch12 · 2026-09-18
- [source] Open-source Benchmark Heaven compares models by cost per actual task, not per-token price — airesearch12 · 2026-09-18
1 near-duplicate retellings: airesearch12