AIRA₂ Research Agents Hit 81.5% on MLE-bench-30, Beating Prior SoTA of 72.7%

mariofilhoml · x · 2026-09-25

Martin Josifoski's team unveiled AIRA₂, a next-generation AI research agent for ML built to remove key scaling bottlenecks:

The sharer, mariofilhoml, highly recommends the "Closing the Generalization Gap" sections and notes the ablations support his view that overfitting the validation set (via model selection after hyperparameter optimization, not literal training on it) is rarely a big issue in practice.

Related event: AIRA₂ Research Agent Tops MLE-bench with 81.5% Win Rate Above Human Baseline(2 posts)→

Original post →

More from Models

Models channel →