Perfect task routing beats the best single model by 15 points in pass@1

ZainHasan6 · x · 2026-07-21

The chart compares a best single model against perfect per-task routing. - Best single model (`Sol`): **72.3%** pass@1 - Perfect per-task router (`oracle`): **86.7%** - Any-of-three union (`pass@4`): **96.5%** The takeaway is that routing each task to its ideal model can add a large performance gain even at the frontier; the post also claims an oracle routing setup across `{Kimi K3, Fable 5, GPT 5.6 Sol}` yields a **+15%** boost over frontier SOTA.

Original post →

More from Research

Research channel →