DRACO benchmark still puts fusion systems ahead of single frontier models
iamtrask · x · 2026-07-22
OpenMined reran an OpenRouter-style experiment on DRACO, a deep-research benchmark, and found that fusion models still dominate.
- The top 5 systems are all fusions, and the best single frontier model only placed 6th.
- The post says this matches observations from four startups and Los Alamos National Lab.
- The chart shows fusion combinations such as Fable + GPT-5.5, Opus + GPT-5.5 + DeepSeek, and DeepSeek + Kimi + GPT-5.5 outperforming solo models by a wide margin.
The author frames this as evidence that combining models remains the best path on deep-research tasks, and that more benchmarks are coming.
Related event: DRACO Benchmark Shows Routing Systems Outperform Single Frontier Models(2 posts)→
More from Research
- Writing Facts into Transformers Without Training: New COLM Paper — HazyResearch · 2026-07-23
- Hazy Research points to a simpler MLP view in pretrained language models — HazyResearch · 2026-07-23
- AI capabilities are improving faster than the way we measure them — ravi_iitm · 2026-07-23
- Fable claims to disprove a century-old conjecture in algebraic geometry — JensHonack · 2026-07-23
- DataFlow-Harness builds editable LLM data pipelines with MCP and code agents — _akhaliq · 2026-07-23
- PeptiVerse adds SOTA models for peptide solubility, toxicity and binding — ahandvanish · 2026-07-23