AREX follow-up links the paper, project page and reported benchmark results
imjustnewatai · x · 2026-07-26
Follow-up post with receipts for AREX: the paper, project page, diagrams, and released model.
It repeats the reported results: AREX-Turbo is a 4B dense model that beat Qwen3.5-35B on 5 of 6 evaluations; AREX-Base has 122B total parameters with 10B active and beat Qwen3.5-397B across all six reported evaluations. The author also reiterates the claimed +22.9 point BrowseComp gain from autonomous context updating plus recursive auditing, while noting the results are from the research team and have not yet been independently replicated.
Related event: Chinese Team Releases AREX, a Recursive Research Agent(2 posts)→
More from Models
- Grok 4.5 is being pitched as a one-person game studio — FinanceYF5 · 2026-07-26
- Yishan says Codex-5.5 and 5.6 are good enough after switching back from Claude — yacineMTB · 2026-07-26
- Gemini portrait misses revive the case for user-adjustable AI “knobs” — yacineMTB · 2026-07-26
- Opus 5 vs Fable 5: same 3D Flappy Bird prompt, 250k vs 130k tokens — TAbrodi · 2026-07-26
- Terra High reasoning impresses, while Grok 4.5 is next on the test list — MickeySteamboat · 2026-07-26
- OpenAI discloses a security incident and Anthropic launches Claude Opus 5 in a packed week — btibor91 · 2026-07-26