Open-source Benchmark Heaven compares models by cost per actual task, not per-token price
airesearch12 · x · 2026-09-18
Developer airesearch12 released Benchmark Heaven, a free open-source project arguing that price-per-million-tokens is a lie: a cheaper model that thinks twice as long isn't actually cheaper. The tool ranks models by cost per actual task, respecting provider filters and accounting for tokens in/out, reasoning, and caching. By its metric, GLM-5.3 is less benchmaxxed than Gemini 3.8 Flash.
More from Models
- Ternary-Bonsai-2-27B, a 2-bit ternary model for on-device inference, trends on Hugging Face — prism-ml · 2026-09-18
- Mystery "Stealth Union Alpha" Model on OpenRouter Baffles Redditors — Iory1998 · 2026-09-18
- Typesafe AI launches general-purpose steerable low-latency classifier — andreisavu · 2026-09-18
- Cactus Releases Needle 3: An 8-29MB Foundation Model Running 4k tokens/s on a Raspberry Pi 5 — airesearch12 · 2026-09-18
- Classifier scores 100 YouTuber videos for sales intent in 12s at ~$0.02 — eptwts · 2026-09-18
- Dan Shipper gets early vibe check of secret new LLM from InstructGPT author Diogo — danshipper · 2026-09-18