SDABench Evaluates LLM Scientific Discovery

HKUST · hf · 2026-07-16

SDABench: An LLM Benchmark for Scientific Discovery Capabilities

This work points out that existing scientific data analysis benchmarks often only evaluate code execution or pipeline completion, ignoring the distinct types of claims inherent in scientific analysis: hypothesis exploration, statistical inference, mechanistic explanation, etc. To address this, the authors propose SDABench, shifting the evaluation focus to six capabilities: description, exploration, inference, prediction, causality, and mechanism.

Benchmark Composition

Key Findings

Additional Contributions

The authors also introduce a five-stage error analysis framework to pinpoint the exact failure modes of LLMs within the scientific analysis chain.

Original post →

More from Research

Research channel →