Benchmark Heaven aggregates 100 benchmarks across 800 models, priced per task instead of per token
airesearch12 · x · 2026-09-21
A new model-overview site called Benchmark Heaven aggregates 100 benchmarks across 800 models. It introduces what the author believes is the first 'Benchmaxxing' composite score, a new Composite Score, and — notably — prices models by cost per task rather than dollars per token, aiming to give a truer picture of real-world value.
More from Models
- Observation: Models Differ Wildly in Tool-Call Intermediate Steps — DanielLockyer · 2026-09-21
- Rumor: Anthropic's Opus 5.5 landing this week, said to beat Astra on price and quality — daniel_mac8 · 2026-09-21
- Burkov calls stealthy LLM-rival project Jev 'BS' over speed and calibration claims — burkov · 2026-09-21
- Moonshot and Tencent Hunyuan both building Flash models to target agent inference costs — TheZachMueller · 2026-09-21
- Dev Launches Made With Jev, a Free Directory Cataloging Demos, Tools and Skills for the New Model — Sea_Supermarket_5891 · 2026-09-21
- Which Model Actually Understands Reverse Engineering? Dev Seeks MCP Workflow for Ghidra and IDA Pro — obese_coder · 2026-09-21