Price per million tokens is a lie: new benchmark scores models by cost per actual task
airesearch12 · x · 2026-09-18
- Per-token pricing misleads: A cheaper model that thinks twice as long isn't actually cheaper.
- Benchmark Heaven scores models by cost per actual task, factoring in provider filters, tokens in/out, reasoning length, and caching — results may surprise you.
- It also introduces a Benchmaxxing score to detect test-set contamination: models that score 99% on widely-quoted benchmarks but falter on new or private ones nobody can train for.
More from Models
- Researchers: RL training has made models' theory of mind 'terrible' — voooooogel · 2026-09-18
- Uncensored Qwen3.8-27B agentic GGUF quant hits Hugging Face trending — cyjin-yl · 2026-09-18
- OpenAI discloses 6 misalignment reports: models hid mistakes, hunted leaked API keys — DynamicWebPaige · 2026-09-18
- Rumor roundup: Grok 4.7 due this week, OpenAI near another Millennium Problem, Google allegedly using RSI — haider1 · 2026-09-18
- Four practical ways to use Jev for agent harnesses: judging, routing, subagent orchestration — omarsar0 · 2026-09-18
- Ternary Bonsai 2: 27B Model Under 6GB Runs In-Browser on WebGPU, Keeps 98.2% Quality — xenovatech · 2026-09-18