AREX, a recursive research agent, reportedly beats larger Qwen3.5 models in deep-research benchmarks
imjustnewatai · x · 2026-07-26
A Chinese research team reportedly released AREX, a recursively self-improving research agent that repeatedly audits claims, preserves verified evidence, and turns weak claims into new targeted searches.
The post says the smallest model is a 4B dense version that reportedly beat Qwen3.5-35B on five of six deep-research benchmarks. The larger AREX-Base uses 122B total parameters with 10B active and reportedly beat Qwen3.5-397B across all six reported evaluations, including the highest reported WideSearch-en score. The thread also highlights the paper’s claimed ablation gains: autonomous context updating adds 11.8 points, recursive auditing adds another 11.1, for a combined +22.9 on BrowseComp. The author stresses this is not full recursive self-improvement of the neural network itself, but iterative improvement of the research state, strategy, and answer generation.
Related event: Chinese Team Releases AREX, a Recursive Research Agent(2 posts)→
More from Models
- Frontier models keep shipping so fast that no single model has a moat — 0xsachi · 2026-07-26
- A client project convinced one builder the agent harness matters more than the model — alexcovo_eth · 2026-07-26
- OpenAI’s GPT-6 and RSI rumors get a source map with internal compute and self-play clues — imjustnewatai · 2026-07-26
- OpenAI’s GPT-6 may be a memory-first system that helps train its successor — imjustnewatai · 2026-07-26
- Comparison post says Fable 5 is faster and cheaper than Opus 5 — bindureddy · 2026-07-26
- Receipts compare Google TPU, Gemini and xAI’s 10T training claims — imjustnewatai · 2026-07-26