Open-source Benchmark Heaven compares models by cost per actual task, not per-token price
airesearch12 · x · 2026-09-18
Developer airesearch12 released Benchmark Heaven, a free open-source hobby project arguing that price-per-million-tokens is misleading: a cheaper model that thinks twice as long isn't actually cheaper. The tool ranks models by cost per actual task, accounting for tokens in/out, reasoning, and caching, with provider filters. By its metric, GLM-5.3 turns out to be less benchmaxxed than Gemini 3.8 Flash.
More from Models
- Cactus Releases Needle 3: An 8-29MB Foundation Model Running 4k tokens/s on a Raspberry Pi 5 — airesearch12 · 2026-09-18
- User flags Claude quietly removing promised "50% extra limits till Sept 20" from pricing page — ThePeterMick · 2026-09-18
- Unverified: Gemini 4 benchmark page spotted, 4 Flash claimed to beat Fable 5.1 — teortaxesTex · 2026-09-18
- DeepSeek v4.1 flash TTFT comparison: Together AI crushes rivals on pre-warmed queries — zhyncs42 · 2026-09-18
- Fine-tuned personas all refuse unsafe requests, and eval scores belong to the harness, not the weights — le_james94 · 2026-09-18
- Scores belong to the system, not the weights: Opus 4.6 jumps 0% to 97.1% with a harness — le_james94 · 2026-09-18