AgentIR: Retrievers That Read Agent Reasoning Hit 68% Accuracy
hllo_wrld · x · 2026-08-01
Two papers from the University of Waterloo's R2L Lab were accepted to COLM:
- AgentIR: Deep Research agents generate explicit natural language reasoning before search calls, but existing retrievers ignore this. The paper introduces a "Reasoning-Aware Retrieval" paradigm that jointly embeds the agent's reasoning trace with its query. Combined with the DR-Synth data synthesis method, their trained AgentIR-4B model achieves 68% accuracy on the BrowseComp-Plus benchmark using the Tongyi-DeepResearch agent. This significantly outperforms conventional embedding models twice its size (50%) and BM25 (37%).
- LakeQuest: A new benchmark designed for answering questions within messy data lakes—environments full of tables, documents, and broken metadata—where every answer must be traced back to concrete evidence, addressing the limitations of clean-text QA benchmarks.
Related event: AgentIR Enables Retrievers to Understand Agent Reasoning(3 posts)→
More from coding & agent
- OpenAI Hits 1 Billion Users, Codex Agents Drive 99.8% of Internal Tokens — firstadopter · 2026-08-01
- Anthropic Agent Accidentally Published Malware to Steal SSH Keys, Researcher Finds — mariofilhoml · 2026-08-01
- Claude Code Costs 3.7x More Than Open-Source Agents in Task Benchmark — Teknium · 2026-08-01
- Grok Build Update: Permanent Session Deletion and Diagnostics — XFreeze · 2026-08-01
- Developer Hits 3,000 GitHub Contributions Using Codex, Consumes 10B Tokens Monthly — DeryaTR_ · 2026-08-01
- DeepSeek-V4-Flash-High Tops Price-Performance in Frontend Code Arena, Ranks #7 Overall — arena · 2026-08-01