AREX, a recursive research agent, reportedly beats larger Qwen3.5 models in deep-research benchmarks

imjustnewatai · x · 2026-07-26

A Chinese research team reportedly released AREX, a recursively self-improving research agent that repeatedly audits claims, preserves verified evidence, and turns weak claims into new targeted searches.

The post says the smallest model is a 4B dense version that reportedly beat Qwen3.5-35B on five of six deep-research benchmarks. The larger AREX-Base uses 122B total parameters with 10B active and reportedly beat Qwen3.5-397B across all six reported evaluations, including the highest reported WideSearch-en score. The thread also highlights the paper’s claimed ablation gains: autonomous context updating adds 11.8 points, recursive auditing adds another 11.1, for a combined +22.9 on BrowseComp. The author stresses this is not full recursive self-improvement of the neural network itself, but iterative improvement of the research state, strategy, and answer generation.

Related event: Chinese Team Releases AREX, a Recursive Research Agent(2 posts)→

Original post →

More from Models

Models channel →