Benchmark Heaven aggregates 264 LLM benchmarks and 870 models with cost-capability analysis
airesearch12 · x · 2026-09-28
A free Beta site called Benchmark Heaven puts 264 LLM benchmarks and 870 models in one place, with detailed cost-vs-capability analysis across providers.
Highlights:
- Multiple composite and per-category rankings (coding agent, agentic tool use, science, long context), with an optional "Benchmaxxing" signal
- Flexible cost bases: fixed I/O blends, output-only, 10:1/30:1/100:1 ratios, or task workloads as measured on OpenRouter
- Regional filters for hosting, provider company, and lab location (China/EU/US), with EU hosting verified per model against provider docs — global deployments and EU billing regions don't count
- Data confidentiality filter flagging providers that train on or retain your prompts
Launched tongue-in-cheek as "Harold filed for forty years, now it's one drawer" — free, which he finds frankly rude.
More from Models
- Does an unguarded Opus still get to be called Opus? — repligate · 2026-09-28
- AI-generated hyperhidrosis treatment plan impresses with trial citations and dosing detail — chaumian · 2026-09-28
- User logs 25 cases of ChatGPT inventing straw-man arguments, then repeating them after apologizing — LAguy8394 · 2026-09-28
- Soap Dispenser Benchmark: five flagship models compete on one HTML animation prompt — Foxiya · 2026-09-28
- Community Speculates OpenAI's New Agent Name: Reviving "Orion" Over "o." — brandon_galang · 2026-09-28
- Princeton professor points out Ember and GLM aren't on the cost-performance Pareto frontier — random_walker · 2026-09-28