Benchmark Heaven Launches: 100 Benchmarks, 800 Models, Real Per-Task Costs
airesearch12 · x · 2026-09-18
Benchmark Heaven (BETA) is a free, ad-free, unfunded AI model leaderboard aggregator:
- 100 benchmarks and 800 models in one place
- 'Benchmaxxing' score — claimed to be the first of its kind
- Real cost per task, not just $ per token
- Composite score plus category rankings (coding, agentic, science, long context)
- Filters for hosting region, data confidentiality, open weights, and customizable I/O price blends and workloads
More from Models
- Cactus Releases Needle 3: An 8-29MB Foundation Model Running 4k tokens/s on a Raspberry Pi 5 — airesearch12 · 2026-09-18
- User flags Claude quietly removing promised "50% extra limits till Sept 20" from pricing page — ThePeterMick · 2026-09-18
- Unverified: Gemini 4 benchmark page spotted, 4 Flash claimed to beat Fable 5.1 — teortaxesTex · 2026-09-18
- DeepSeek v4.1 flash TTFT comparison: Together AI crushes rivals on pre-warmed queries — zhyncs42 · 2026-09-18
- Fine-tuned personas all refuse unsafe requests, and eval scores belong to the harness, not the weights — le_james94 · 2026-09-18
- Scores belong to the system, not the weights: Opus 4.6 jumps 0% to 97.1% with a harness — le_james94 · 2026-09-18