Open-source Benchmark Heaven compares models by cost per actual task, not per-token price

airesearch12 · x · 2026-09-18

Developer airesearch12 released Benchmark Heaven, a free open-source project arguing that price-per-million-tokens is a lie: a cheaper model that thinks twice as long isn't actually cheaper. The tool ranks models by cost per actual task, respecting provider filters and accounting for tokens in/out, reasoning, and caching. By its metric, GLM-5.3 is less benchmaxxed than Gemini 3.8 Flash.

Related event: Benchmark Heaven Launches: Aggregates 100 Benchmarks, 800 Models with Real Task-Cost Comparison(8 posts)→

Original post →

More from Models

Models channel →