DiscoverPhysics: 22 tweaked-physics worlds stump top LLM agents, only half solved
burny_tech · x · 2026-09-26
Can LLMs discover new laws of physics, or do they just recall established science? DiscoverPhysics, a new interactive benchmark on arXiv, tackles exactly this.
- It generates 22 simulated N-body worlds whose physics deliberately deviates from ours: screened and fractional-power gravity, multi-species couplings, hidden dark-matter-like particles, time-varying interactions, and more
- Each agent must propose multiple rounds of experiments, observe raw trajectory data, then submit a natural-language explanation of the world's physics plus a Python implementation of the inferred law
- Evaluation runs on two axes: trajectory MSE on held-out particles, and an LLM-judged explanation score against an expert-written rubric
- Headline finding: across 11 frontier models, the strongest agents pass only about half of the worlds
Since solving a world requires designing informative experiments and revising hypotheses over a full experimental history, the benchmark probes long-horizon scientific reasoning that pure recall cannot fake. The team includes Andrew Gordon Wilson, Pavel Izmailov and other well-known researchers.
More from Research
- RGBD20K: new RGB-D segmentation benchmark with 20K image pairs and 160 categories — UNT · 2026-09-26
- Formally verified Ethereum consensus proposal moves toward 4-8x faster finality — DavideCrapis · 2026-09-26
- Fewer than 10% of preregistered studies include multiple-hypothesis corrections — RexDouglass · 2026-09-26
- Reproducibility study finds 63% of reproduced papers had at least one coding error — RexDouglass · 2026-09-26
- Only 28.4% of papers fully reproducible from raw data, replication study finds — RexDouglass · 2026-09-26
- 67 Teams Replicated 64 Psychology Papers, Only ~44% Held Up — RexDouglass · 2026-09-26