EMNLP 2026 paper: only 7.1% citation recall in deep research reports, new algorithm traces errors to agents
mohitban47 · x · 2026-09-23
An EMNLP 2026 main-conference paper finds that many sentences in deep research reports aren't supported by their citations: recall is 58.7% on NVIDIA AI-Q and just 7.1% on TrajectoryKit. The authors propose an algorithm that traces each error to the specific agent that introduced it and classifies the error type.
Related event: EMNLP Paper Finds Deep Research Reports' Citation Support as Low as 7.1%(2 posts)→
More from Research
- Microsoft's Agensh Scales Multi-Agent Systems to 1,024 Agents Without a Central Orchestrator, Boosting Test-Pass Rate to 55% — andrew_n_carr · 2026-09-23
- SemiAnalysis: Engram DRAM offloading delivers up to 50% better perf, upstreamed to vLLM — bookwormengr · 2026-09-23
- ValsAI says 10 Claude Opus 5.5 agents built a Lean-verified algorithm improvement in 15 hours — alejandroll10 · 2026-09-23
- Stanford-led Terminal-Bench-Science debuts: top model solves just 30% of 70 expert tasks — geoffwolfe · 2026-09-23
- Untrained Qwen beats purpose-trained decision model by 35 points in real repo test — KingPinX · 2026-09-23
- UCL study finds LLMs share a Fisher-Rao information geometry that enables low-distortion model interventions — UniversityCollegeLondon · 2026-09-23