Frontier labs still trail routed model systems on DRACO deep-research tasks
iamtrask · x · 2026-07-22
Andrew Trask says frontier AI labs have lost the “deep research” frontier on DRACO, a benchmark for deep-research tasks.
- Since May, OpenRouter, Sakana AI, and Los Alamos National Laboratory have each reported that routed combinations of models beat individual frontier systems.
- OpenMined re-ran the OpenRouter experiment and says the result still holds a month later.
- In one example, a routed system reached Fable-level quality at roughly half the price of Anthropic’s Fable API, and the top single frontier model only placed sixth in the reproduced setup.
Related event: DRACO Benchmark Shows Routing Systems Outperform Single Frontier Models(2 posts)→
More from Research
- Programmable Cellular Automata: CA rules as readable code for explainability — Amidos2006 · 2026-09-11
- GEVIBench launches as a comprehensive benchmark for comparing voltage indicators — drmichaellevin · 2026-09-11
- Gaussian Light Transport: 13D Gaussian Mixtures Speed Up Global Illumination — ssh4net · 2026-09-11
- Llama Loves Pirates — Goodfire's Tom McGrath on teaching math without the pirate style — Machine Learning Street Talk · 2026-09-11
- Fortnow: P vs NP beyond AI's reach, but NP vs L separations could fall — fortnow · 2026-09-11
- MaP-WAM tackles non-Markovian robot manipulation with memory-grounded planning — Sizhe Zhao · 2026-09-11