BAAI's AREX-2 Trains Self-Improving Agents, Hits 92.2 on GAIA and 81.8 on MLE-bench Lite
BAAI · hf · 2026-10-01
BAAI presents AREX-2, which trains test-time self-improvement by synthesizing long-horizon reflective trajectories from ML and algorithmic programming tasks. The Qwen3.8-27B-based agent scores 81.8 on MLE-bench Lite, 70.7 on Frontier-CS, and transfers to deep research with 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA, improving as its round budget grows.
More from coding & agent
- Hybris MCP Server lets AI assistants manage SAP Commerce Cloud instances — modelcontextprotocol · 2026-10-01
- cua-speedrun: CMU benchmark shows 4.4x speed gap between equal-scoring computer-use agents — arankomatsuzaki · 2026-10-01
- M-Anchor: a deterministic, zero-LLM Python gate that blocks unsupported LLM record updates — Informal-Winter-3190 · 2026-10-01
- Dev Burns Through Claude Weekly Quota in Under Two Days on Opus 4.5 Alone — yihui_indie · 2026-10-01
- Split coding workflow: Opus 5.5 xhigh for planning, Sol 6.1 high for the bulk of coding — haider1 · 2026-10-01
- FV expert: AI folks' view of formal verification is 15 years out of date — tianyin_xu · 2026-10-01