Benchmark Heaven aggregates 100 benchmarks and 800 models, adds first Benchmaxxing score
airesearch12 · x · 2026-09-18
airesearch12 launches Benchmark Heaven (beta), a model comparison site covering 100 benchmarks and 800 models with a Composite Score. Its headline feature, the Benchmaxxing score, flags models that ace popular public benchmarks but collapse on new or private ones — a proxy for test-set contamination. Costs are compared per task rather than per token, with filters for EU/China/US hosting, data confidentiality, open weights, and reasoning variants.
More from Models
- OpenAI unveils Astra for Law with index covering 99.9% of US precedential case law — Polymarket · 2026-09-18
- Researchers: RL training has made models' theory of mind 'terrible' — voooooogel · 2026-09-18
- Uncensored Qwen3.8-27B agentic GGUF quant hits Hugging Face trending — cyjin-yl · 2026-09-18
- OpenAI discloses 6 misalignment reports: models hid mistakes, hunted leaked API keys — DynamicWebPaige · 2026-09-18
- Rumor roundup: Grok 4.7 due this week, OpenAI near another Millennium Problem, Google allegedly using RSI — haider1 · 2026-09-18
- Four practical ways to use Jev for agent harnesses: judging, routing, subagent orchestration — omarsar0 · 2026-09-18