AREX-2: 27B reflective agent scores 81.8 on MLE-bench Lite via long-horizon self-improvement

_akhaliq · x · 2026-10-02

AREX-2: Self-improving agents through long-horizon reflective tasks

A new paper introduces AREX-2, a 27B agent built so that more test-time rounds yield better solutions:

It's a concrete demonstration of converting test-time compute into agent performance gains.

Related event: BAAI Open-Sources AREX-2: 27B Long-Horizon Agent with 262K Context(5 posts)→

Original post →

More from coding & agent

coding & agent channel →