BAAI's AREX-2 Trains Self-Improving Agents, Hits 92.2 on GAIA and 81.8 on MLE-bench Lite

BAAI · hf · 2026-10-01

BAAI presents AREX-2, which trains test-time self-improvement by synthesizing long-horizon reflective trajectories from ML and algorithmic programming tasks. The Qwen3.8-27B-based agent scores 81.8 on MLE-bench Lite, 70.7 on Frontier-CS, and transfers to deep research with 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA, improving as its round budget grows.

Original post →

More from coding & agent

coding & agent channel →