Paper Reveals Deep Research Vulnerability: Misleading Info Triggers False Conclusions

Pengyu Zhu · hf · 2026-07-31

This paper investigates the reliability of Deep Research agents in open information environments. The study reveals that plausible but misleading knowledge can propagate through the agent's workflow and ultimately result in false conclusions within the final report.

To systematically study this failure mode, the authors propose the MisKnow-Agent framework, generating a dataset of 5,933 quality-controlled misleading instances based on the DeepResearch Benchmark.

Experiments demonstrate that even limited exposure to misleading information causes both open-source and closed-source agents to adopt false conclusions. While search-enabled verifier models can identify these instances during focused validation, they still get misled during long-horizon research. The authors evaluated various pre- and post-research defenses, finding that they only mitigate but fail to completely prevent false-conclusion adoption. The findings suggest that reliable Deep Research requires robust evidence verification and correction capabilities at both the model and framework levels.

Original post →

More from Safety

Safety channel →