DDR-Bench reveals LLM 'long-term exploration amnesia'; Agent frameworks alone are insufficient
青稞AI · wechat · 2026-08-24
Core Problem
Most LLM evaluations focus on "executive intelligence" (following instructions), overlooking the "research intelligence" needed for data science: autonomous goal setting and insight mining. This paper proposes the "Deep Data Research (DDR)" task and the DDR-Bench benchmark.
Methodological Innovation
In the DDR task, models are placed in raw databases without preset questions and must perform exploratory analysis. To solve the evaluation challenge for open-ended tasks, a "checklist" mechanism is introduced, objectively scoring models based on how many preset key facts and hidden correlations their reports uncover.
Key Findings
Experiments reveal a "long-term exploration amnesia" in frontier models:
- Goal drift and context forgetting occur after multi-round tool use;
- Exploration logic remains shallow, lacking systematic deep-digging;
- Models tend to output "generic hallucinations" rather than real data-driven findings.
Conclusion & Suggestions
The study confirms that external Agent frameworks alone cannot solve these bottlenecks. Achieving true "research intelligence" requires upgrading intrinsic model mechanisms, such as introducing curiosity-driven reinforcement learning. The benchmark's evaluation paradigm is highly valuable and recommended for reproduction on latest flagship models.
Paper Link
More from Research
- CfP: MRL 2026 Workshop Co-located with EMNLP — davlanade · 2026-08-25
- AI Fact-Checker Audit: 1 in 18 Citations Were Fabricated — jonathancheckwise · 2026-08-25
- Study Shows Long-Form Context Induces Internal Drift Bypassing Safety — PresentSituation8736 · 2026-08-25
- Chinchilla-style scaling laws found for human motion: the fifth scalable modality — andrew_n_carr · 2026-08-25
- Using AI Pipelines to Process Hebrew Memory Books: From Cleaning to Knowledge Graphs — aloncarmel · 2026-08-25
- MIT Study: Aging Brains Maintain Language Networks Like LLMs Trained on Lifetime Data — MacrinePhD · 2026-08-25