AtomWorld: LLMs Struggle with Atomic Operations

新智元 · wechat · 2026-07-14

This article introduces AtomWorld, a new materials science benchmark designed to evaluate large models on "atomic-level spatial operations" rather than just text comprehension.

Key conclusions include:

AtomWorld's evaluation pipeline auto-generates "structure-text" paired samples and uses structure matching tools to compare model outputs against ground truths, quantifying whether the model correctly altered the atomic arrangement per instructions. The article notes that while external tools can help output precise coordinates, they cannot replace the model's own decision-making regarding where atoms belong or which atoms form the target region.

The author broadens this issue to general scientific agents: in fields like materials modeling, molecular design, and automated experiments, AI must reliably execute actions, not just interpret knowledge, to truly participate in scientific research.

Original post →

More from Apps

Apps channel →