AREX-2: 27B reflective agent scores 81.8 on MLE-bench Lite via long-horizon self-improvement
_akhaliq · x · 2026-10-02
AREX-2: Self-improving agents through long-horizon reflective tasks
A new paper introduces AREX-2, a 27B agent built so that more test-time rounds yield better solutions:
- Trained on verifiable ML and algorithmic tasks using long-horizon reflective tasks
- Scores 81.8 on MLE-bench Lite and 70.7 on Frontier-CS
- Capabilities transfer to deep research scenarios
It's a concrete demonstration of converting test-time compute into agent performance gains.
Related event: BAAI Open-Sources AREX-2: 27B Long-Horizon Agent with 262K Context(5 posts)→
More from coding & agent
- SerenityOS Creator Maps Open Source's Five Stages of Grief Over AI Code — vivekhaldar · 2026-10-02
- Solo Founder Launches OpenSwitchboard, a Matchmaker for AI Assistants Handling Real-World Errands — EnvironmentalRice348 · 2026-10-02
- App idea: an MCP-powered alarm that wakes you when your agent finishes — Angaisb_ · 2026-10-02
- Wellness Project Launches MCP Connector Bringing Health Data Into Claude — turnnoblindeye · 2026-10-02
- How an FDE maps business processes for AI deployment using three sources — vasuman · 2026-10-02
- Rhyven Marketplace: Open-Source MCP-Based App Store for AI Agents Launches Linux Preview — Due-Telephone8276 · 2026-10-02