Martian says routing across 44 LLMs cuts errors 46% vs best single model on 16 benchmarks
Arindam_1729 · x · 2026-09-04
Martian's AI Frontier argues that standard benchmarks—measuring a single model on a single run—systematically underestimate what AI can achieve.
Key points:
- By routing requests across 44 LLMs, they build a "Capability Frontier": the best possible performance at every cost level
- Across the 16 most-used benchmarks (TerminalBench, LiveCodeBench, etc.), this yields 46% fewer errors than the single best LLM, or cost savings at matched SOTA quality
- Takeaway: no single model dominates every task; each benchmark misses the majority of model capabilities
Includes an interactive site and an academic paper.
More from Models
- GPT-6 Astra Tops ValsAI Code Migration Benchmark at 68% Accuracy, 2-4x Faster — sandersted · 2026-09-04
- OpenAI Launches GPT-6 Astra: An Agent That Can Do Anything on Your Computer — astralmatrix · 2026-09-04
- Andrew Ng: We need harder evals — have frontier models chat with me — andrewgwils · 2026-09-04
- Epoch AI launches FrontierMath Erdős benchmark of 68 unsolved problems; Astra tops it at 2/68 — keviv9 · 2026-09-04
- OpenAI Launches GPT-6 Astra: 99.9% on ARC-AGI-3, but Independent Evals Call It Uneven — Latent Space · 2026-09-04
- Liquid AI launches Nanos: task-specific 350M-2.6B models that run on-device — JosephJacks_ · 2026-09-04