Can AI Follow In Einstein's Footsteps?
Michael Shalyt, Nathan Regev, Marin Soljačić, Ido Kaminer
physics.hist-ph, cs.AI, physics.pop-ph
2026-07-30
A Technion/MIT Perspective: AI for physics runs in reverse, from symbolic equations to black-box predictors like AlphaFold, with zero principle-level theories so far.
AlphaFold solves protein structures, GraphCast beats numerical weather solvers, and AI keeps notching wins across the physical sciences. This Perspective from Technion and MIT (Shalyt, Regev, Soljačić, Kaminer) asks a more uncomfortable question than "how accurate can it get": why has AI still not produced a single principle-level theory to set beside general relativity or the Standard Model? Their diagnosis is that the center of gravity of AI for physics is drifting in the opposite direction from the history of human physics.
Human physics, in broad strokes, climbed from pattern prediction (Babylonian cycles, Mayan astronomy) through phenomenological laws (Kepler, Planck's black-body formula) to principle-based theories (Newton, Maxwell, Einstein), with the level of "understanding" rising throughout. AI's contributions have moved the other way: early symbolic regression (BACON, Eureqa, SINDy, AI Feynman) still mined explicit equations from data, while today's marquee systems discard equations and act as high-accuracy predictors.
The paper's main analytical device is a three-way taxonomy that sorts discoveries by their level of mathematical abstraction:
The categories are a ruler for the size of the epistemic leap, not mutually exclusive labels. The verdict is blunt: AI has impressive outputs in A and B, and the C column is empty. In the authors' own words, no AI has made a discovery of Category C.
Why is C so hard? They split reasoning into three modes. Induction fits hypotheses to data; deduction computes consequences from a given framework; AI is good at both. The bottleneck is the third mode, abduction: guess a set of provisional axioms from abstract principles (symmetry, simplicity), then derive falsifiable consequences to test. Einstein did not build general relativity by fitting Mercury's perihelion. He adopted symmetry constraints (equivalence, covariance), derived a candidate theory, and only then sought empirical verification. The paper argues current AI has not reproduced this "principles first, verification later" mode.
The deeper obstacle is not novelty. In a large enough search space, novelty is cheap. What is hard is judging the payoff-per-effort of each novel idea: whether it can be executed, whether it will matter, whether it yields falsifiable tests. This judgment is the "gut feeling" of top physicists, rarely written down as a procedure and therefore almost absent from training data.
This is a Perspective, not an experimental paper; there are no benchmarks. Its "results" are the taxonomy above and the repeated observation that AI has delivered A- and B-class work while producing zero C-class discoveries. The table mirrors the paper's central classification:
| Category | Type | Human examples | Has AI done it? |
| A new capability | prediction, device, application | ballistics, transistors, lasers, collider phenomenology | yes: AlphaFold, GraphCast, PINN surrogates |
| B new equation | theory, ansatz, formula | Kepler's laws, Lorentz force, black-body formula | yes: AI Feynman recovered equations, symbolic regression gave a dark-matter formula, LLMs derive scattering amplitudes |
| C new framework | new math language, symmetry principle | Newton's calculus, Maxwell's field theory, operator QM, gauge theory | zero; no AI has made a Category C discovery |
One piece of supporting evidence: the authors see no fundamental ceiling on AI creativity, and cite cases where AI went beyond prior knowledge to make surprising choices, including AlphaGo's Move 37, an OpenAI internal model disproving Erdős's planar unit-distance conjecture, and Anthropic's Fable producing a counterexample to the Jacobian conjecture. What is missing is not creativity but the principle-guided mode of discovery.
For anyone building AI for science, this is a sober strategic map. It separates "predicting accurately" from "producing new theory," and warns against reading AlphaFold-style wins as proof that the rest is just engineering scale. Keep tilting toward black-box prediction and AI will get ever better at predicting within known frameworks while possibly never proposing a contender to quantum gravity.
A few routes are concrete enough to act on. Embedding explicit physical laws as inductive bias lets models train on less data and generalize better. Wiring computer algebra systems (xAct, Cadabra, FORM, FeynCalc) into agent loops creates a "theory-building sandbox" where a proposed symmetry is checked for consistency before it ever sees data. The effective-field-theory community already automates EFT matching (matchmakereft, Matchete) and even inverts the workflow to search high-energy theories from a low-energy Lagrangian.
Honestly, this reads more as a research agenda than a validated plan. It points a direction and offers no guarantee.
The authors leave several exits open.
First, the complexity trap. Not every domain hides an elegant equation underneath. In condensed matter, chemistry, and biology, contributions at different scales often cannot be separated, and accurate prediction without transparent explanation may be the best achievable outcome. The paper explicitly asks of AlphaFold whether the absence of a closed-form folding formula is intrinsic or whether we have missed a structural law underneath. It gives no answer.
Second, a provocative possibility: AI may discover new mathematics humans cannot parse, communicated at a higher level between AI agents, where forcing human understanding only slows progress. Emergent communication and continuous-latent reasoning already hint at this. The authors flag it as "Ipcha Mistabra" (Aramaic for "the opposite is unproven") and admit it is not ruled out.
Third, the most basic doubt: maybe the bottleneck is physics itself, not AI. Maybe the era of simple, experimentally testable theoretical revolutions is over for humans as well. The authors concede the past few decades can be read that way, but insist it is too early to declare theoretical physics finished.
One weakness stands out on its own. The spine of the argument is the "reverse trajectory," yet the authors repeatedly stress that it is only a rough sketch rather than a strict chronology, and offer counterexamples such as Ptolemy and Kepler considering their own work principle-based, or Maxwell using mechanical scaffolding he later discarded. Since the spine is an approximate narrative, the conclusion that Category C must remain empty carries a narrative flavor rather than the force of a mechanistic proof.