MBZUAI x Ant Group open-source OPD work: IER selector helps thinking mode too
jiank_uiuc · x · 2026-09-23
The author confirms this is a formal MBZUAI–Ant Group collaboration, credits the team, and shares extra results: adding the IER-based token selector improves math performance in both thinking-on and thinking-off modes, with the strongest gains at 0.1% or 1% token budgets. The linked arXiv page (2609.24432) gives the full abstract: sparse OPD suffers noisy updates from useful-but-single-token teacher guidance; IER, based on a signal-to-noise decomposition, characterizes relative gradient-estimation error, and combining it with usefulness scores matches or exceeds full OPD at 0.1%–1% budgets on math and medical reasoning tasks. Code is open-sourced.
More from Research
- CAIS benchmark: GPT-6 Astra automates 20.8% of remote work, up from 2.5% a year ago — scaling01 · 2026-09-23
- Eval scores reported to 5 significant digits despite 2-5 point error bars, researcher flags — dfrsrchtwts · 2026-09-23
- SkillSpec: intent-masked specification reasoning to catch semantic defects in agent skills — Buaa1 · 2026-09-23
- LLMs start overconfident, then swing underconfident when criticized, Nature MI paper finds — ValerioCapraro · 2026-09-23
- Nature MI paper: LLMs start overconfident, then swing underconfident when criticized — ValerioCapraro · 2026-09-23
- UK AISI paper: fixed-budget evals increasingly understate frontier LLM capability — evijit · 2026-09-23