BAAI's AREX research agent checks answers requirement-by-requirement, hits 82.5% BrowseComp
DeepLearningAI · x · 2026-10-09
Researchers at the Beijing Academy of Artificial Intelligence introduced AREX, a research agent and harness that turns answer checking into an iterative improvement loop: it verifies a tentative answer requirement by requirement, keeps the verified parts, and launches targeted follow-up searches only for the gaps.
- Built a dataset with the harness and fine-tuned two models to work with it: Qwen3.5-4B and the larger MoE Qwen3.5-122B-A10B.
- Scores 82.5% on BrowseComp and 82.0 F1 on WideSearch-en; the harness alone accounts for up to +10 points.
- The fine-tuned 4B model beats the untuned 35B model on 5 of 6 benchmarks, showing harness design plus targeted fine-tuning matters more than raw scale.
More from coding & agent
- Swapping harness lifts GPT-5.6 repo migration from 6.5% to 31%, paper finds — omarsar0 · 2026-10-09
- He replaced a 17-step agent prompt checklist with a 1976 Unix Makefile — 140+ tickets without drift — RandalSchwartz · 2026-10-09
- Developer builds wildfire analysis tool with Opus 5.5: fire progression and hot spot tracking — natesiggard · 2026-10-09
- Subagent-as-a-service: betting on 'expert' agents like Harvey as tool calls for vertical AI — nbaschez · 2026-10-09
- WebMCP Debate: LLM Agents Skip Ads, Putting Ad-Based Business Models at Risk — ConnectRub4818 · 2026-10-09
- At first Grok Bot Meetup, audience-voted idea becomes a 139-marker map app in ~10 minutes — pswider · 2026-10-09