Martian's model router cuts errors 46% vs best single LLM across 16 benchmarks, ICLR oral

minchoi · x · 2026-09-04

Martian shows that routing each task to the right model beats locking into one "best" LLM: 46% fewer errors than the single best model across 16 popular benchmarks (TerminalBench, LiveCodeBench, etc.), backed by an ICLR 2026 oral. Key insight: quoted token pricing is a poor predictor of real cost—some models charge more and think less, others charge less and burn tokens. Every benchmark misses the majority of each model's capabilities.

Related event: Martian's AI Frontier: Model Routing Cuts Error Rates at Same Cost(6 posts)→

Original post →

More from Models

Models channel →