AIRA₂ Research Agents Hit 81.5% on MLE-bench-30, Beating Prior SoTA of 72.7%
mariofilhoml · x · 2026-09-25
Martin Josifoski's team unveiled AIRA₂, a next-generation AI research agent for ML built to remove key scaling bottlenecks:
- State-of-the-art 81.5% on real-world MLE-bench-30 tasks, up from 72.7%;
- Exceeds human SoTA on 6 of 20 diverse AI research tasks in AIRS-Bench (and "hacks" 5 more);
- Shows strong, predictable scaling properties, with lessons already feeding the next iteration.
The sharer, mariofilhoml, highly recommends the "Closing the Generalization Gap" sections and notes the ablations support his view that overfitting the validation set (via model selection after hyperparameter optimization, not literal training on it) is rarely a big issue in practice.
More from Models
- Student Pays $30/Month for Google AI Pro, Still Finds Gemini Unreliable and Hallucinatory — makeitbumthem · 2026-09-25
- Meta's Muse app hit 1.8M iOS downloads in 12 days, beating ChatGPT's 1.3M — Beth_Kindig · 2026-09-25
- Opus 5.5's token usage is far more reasonable: 13-hour session on a 20x Max plan — majidmanzarpour · 2026-09-25
- TypeSafe's Jev: A Model That Can't Write but Decides, Claiming 200x Speed and 400x Cost Savings — rseroter · 2026-09-25
- Prime Intellect pitches highest-throughput GLM 5.3 inference with OpenAI-compatible eval API — willcb · 2026-09-25
- Dev disassembles Claude Code: the "thinking" status labels are just a 15-second timer — ryunuck · 2026-09-25