Meta AI's MIRA splits research agents to fix long-horizon decision learning
dair_ai · x · 2026-10-06
dairai highlights a Meta AI paper on MIRA, addressing the core difficulty of long-horizon research agents: choosing what to investigate next is hard to learn because such decisions are rare in long traces and their effects appear several steps later.
- Architecture: MIRA splits the agent in two — an outer meta-reasoner reads a persistent research record and writes a work order for the next investigation, while a fresh executor carries out each work order.
- Training: Decisions happen only at work-order boundaries, where a critic is trained to forecast remaining return, yielding a single actor-critic (MIRA-AC) that both values partial progress and picks the next investigation.
- Results: Even without training, the split improves theorem proving and open-ended architecture research; with training on the model's own proxy signals, MIRA-AC improves gold scores across all four autoresearch benchmarks.
More from AGI Musings
- Ex-UK energy official warns AI needs 500GW by 2035 and the industry isn't ready — ShakeelHashim · 2026-10-06
- John Langford's COLM 2026 Keynote 'Research at AI Speed' Examines the Industrialization of R&D — JohnCLangford · 2026-10-06
- Critic Calls Out 'Research Taste Doubles Every 3 Months' Eval as Merely Metric Optimization — dhadfieldmenell · 2026-10-06
- Nearly 4 years of ChatGPT: zero direct AI deaths vs 225M from disease and aging — DeryaTR_ · 2026-10-06
- 2048 Satire: Every Job Automated Except Healthcare Admin and Unionized Work — rickasaurus · 2026-10-06
- MIT economist: even the most optimistic AI scenario means major labor disruption — soumitrashukla9 · 2026-10-06