EMNLP paper traces citation errors in deep research systems: the orchestrator causes most
EliasEskin · x · 2026-09-23
An EMNLP 2026 main-conference paper diagnoses citation errors in deep research systems, where many sentences lack supporting citations — recall is 58.7% in NVIDIA AI-Q and just 7.1% in TrajectoryKit. Information gets corrupted between agents like a game of telephone; the authors propose an algorithm that traces each error to the responsible agent and classifies its type. Surprisingly, the orchestrator causes most final-report errors despite having one of the lowest per-agent error rates, and a one-sentence prompt fix raises citation recall by 5%.
Related event: EMNLP Paper Finds Deep Research Reports' Citation Support as Low as 7.1%(2 posts)→
More from coding & agent
- Alibaba contributes $3M to Omarchy to build an ideal agentic OS — AIFlow_ML · 2026-09-23
- Jev Founder Publishes Coding Agent Harness Guide Claiming 200x Speed, 400x Cost Cut — _AustinCalvert_ · 2026-09-23
- AI trading agent backtest: 1,155 decisions in 77 days for $0.09, still trails buy-and-hold — nikola_mr64990 · 2026-09-23
- CoVeR Cuts 62-68% of Agentic Retrieval Verifier Calls Without Losing Accuracy — _reachsumit · 2026-09-23
- Decision Desk turns one support ticket into four typed decisions with confidence and cost — airesearch12 · 2026-09-23
- Successful Agent Retries Can Make Traces More Misleading, Not Less — Sensitive-Parsnip-12 · 2026-09-23