DR Tulu: open 8B deep-research model with evolving rubrics matches OpenAI DR
AkariAsai · x · 2026-09-05
Akari Asai highlights two open research efforts:
- DR Tulu (ICML oral): introduces RLER, reinforcement learning with evolving rubrics that co-evolve with the policy model during training, incorporating newly explored search results and contrasting responses. The resulting DR Tulu-8B is the first fully open model trained for open-ended long-form deep research, beating Tongyi DR by 15.6% on average across four benchmarks and edging out OpenAI DR by 0.7%.
- AgentIR: to be presented at COLM next month.
Author list includes Luke Zettlemoyer, Hannaneh Hajishirzi, Pang Wei Koh and others from UW/AI2.
More from Research
- Tandem Training: RL method makes strong models' reasoning followable by weaker models — erichorvitz · 2026-09-05
- Artificial Analysis discloses full benchmarking methodology spanning a dozen evals — ArtificialAnlys · 2026-09-05
- Artificial Analysis Launches Intelligence Index v4.2 With 40% Private Test Sets to Block Benchmark Gaming — ArtificialAnlys · 2026-09-05
- PolyU's OmniColor unifies multi-signal lineart colorization in an ECCV 2026 paper — jiqizhixin · 2026-09-05
- MIT's SwarmWorld paper: agent swarms win when discoveries accumulate, not when agents get smarter — rohanpaul_ai · 2026-09-05
- Pedro Domingos Proposes Tensor Logic, a Language Unifying Neural and Symbolic AI — pmddomingos · 2026-09-05