DDR-Bench reveals LLM 'long-term exploration amnesia'; Agent frameworks alone are insufficient

青稞AI · wechat · 2026-08-24

Core Problem

Most LLM evaluations focus on "executive intelligence" (following instructions), overlooking the "research intelligence" needed for data science: autonomous goal setting and insight mining. This paper proposes the "Deep Data Research (DDR)" task and the DDR-Bench benchmark.

Methodological Innovation

In the DDR task, models are placed in raw databases without preset questions and must perform exploratory analysis. To solve the evaluation challenge for open-ended tasks, a "checklist" mechanism is introduced, objectively scoring models based on how many preset key facts and hidden correlations their reports uncover.

Key Findings

Experiments reveal a "long-term exploration amnesia" in frontier models:

Conclusion & Suggestions

The study confirms that external Agent frameworks alone cannot solve these bottlenecks. Achieving true "research intelligence" requires upgrading intrinsic model mechanisms, such as introducing curiosity-driven reinforcement learning. The benchmark's evaluation paradigm is highly valuable and recommended for reproduction on latest flagship models.

Paper Link

Original post →

More from Research

Research channel →