Frontier labs still trail routed model systems on DRACO deep-research tasks
iamtrask · x · 2026-07-22
Andrew Trask says frontier AI labs have lost the “deep research” frontier on DRACO, a benchmark for deep-research tasks.
- Since May, OpenRouter, Sakana AI, and Los Alamos National Laboratory have each reported that routed combinations of models beat individual frontier systems.
- OpenMined re-ran the OpenRouter experiment and says the result still holds a month later.
- In one example, a routed system reached Fable-level quality at roughly half the price of Anthropic’s Fable API, and the top single frontier model only placed sixth in the reproduced setup.
Related event: DRACO Benchmark Shows Routing Systems Outperform Single Frontier Models(2 posts)→
More from Research
- Writing Facts into Transformers Without Training: New COLM Paper — HazyResearch · 2026-07-23
- Hazy Research points to a simpler MLP view in pretrained language models — HazyResearch · 2026-07-23
- AI capabilities are improving faster than the way we measure them — ravi_iitm · 2026-07-23
- Fable claims to disprove a century-old conjecture in algebraic geometry — JensHonack · 2026-07-23
- DataFlow-Harness builds editable LLM data pipelines with MCP and code agents — _akhaliq · 2026-07-23
- PeptiVerse adds SOTA models for peptide solubility, toxicity and binding — ahandvanish · 2026-07-23