SDABench Evaluates LLM Scientific Discovery
HKUST · hf · 2026-07-16
SDABench: An LLM Benchmark for Scientific Discovery Capabilities
This work points out that existing scientific data analysis benchmarks often only evaluate code execution or pipeline completion, ignoring the distinct types of claims inherent in scientific analysis: hypothesis exploration, statistical inference, mechanistic explanation, etc. To address this, the authors propose SDABench, shifting the evaluation focus to six capabilities: description, exploration, inference, prediction, causality, and mechanism.
Benchmark Composition
- Covers 5 domains: biology, chemistry, environment, geography, and physics.
- Contains 527 real data samples (SDA-Real) and 6000 synthetic samples (SDA-Synth).
- Provides both multiple-choice and open-ended formats for each sample.
- Constructed via an automated pipeline.
Key Findings
- Evaluated 15 representative LLMs.
- Models perform relatively well on descriptive analysis.
- Performance drops significantly once hypothesis selection, latent process modeling, or mechanistic reasoning are involved.
- Stronger models are better at identifying relevant scopes and variables, but still frequently fail in the following steps:
- Selecting appropriate analysis methods.
- Establishing variable relationships.
- Drawing valid conclusions.
Additional Contributions
The authors also introduce a five-stage error analysis framework to pinpoint the exact failure modes of LLMs within the scientific analysis chain.
More from Research
- OpenAI-style autonomous researchers could become real scientific collaborators — Promptmethus · 2026-07-21
- Soft Clamp cuts tool-call overuse in multi-teacher distillation, from 13.7% to 9.0% — antgroup · 2026-07-21
- ShotPlan adds learnable planning tokens for cinematic multi-shot video generation — Tele-AI · 2026-07-21
- A silicon photonic reservoir chip compensates fiber distortion in real time at 28 Gbps — bravo_abad · 2026-07-21
- A developer maps out six design rules for CLIs that humans and AI agents can both use — yujiezha · 2026-07-21
- GPT 5.6 vs. Claude Fable tested in Dyad AI for Physical AI model tuning — ChrisRackauckas · 2026-07-21