2026-08-24
The CMU Gomes group argues trust in autonomous experiments comes from a re-openable, auditable, correctable record, not from opening the model.
Large language models already plan experiments, drive instruments, and issue pass/fail verdicts. The same group showed this in 2023 with Coscientist in Nature: a natural-language instruction was enough for a model to plan and run chemistry on automated hardware. The usual response is that a system you cannot inspect internally cannot be trusted to produce scientific claims.
This Nature Computational Science Comment from Carnegie Mellon's Gomes group (MacKnight, Novitskiy, Radadiya, Gomes) rejects that reflex. Scientific trust has never required a view into a scientist's head. It requires a record of what was reasoned, done, and measured that can be reopened, audited, and corrected when it is wrong. Interpretability still helps. The gateway is the record.
The four-page body sits behind a paywall. What follows is reconstructed from the public abstract, Figure 1, and the reference list, not from a line-by-line reading of the prose. Field schemas and case write-ups that do not appear in those three sources are not filled in here.
This is a position piece, not a new model or benchmark. The argument is compressed into a four-panel figure.
Figure 1a splits three capabilities. Autonomy can act (a robot arm). Interpretability may explain (a magnifying glass over a network graph). A record reconstructs, audits, and corrects (a verdict with a red cross). Only the third line reaches scientific trust. A robot that runs reactions, and a close look at the model, both fall short.
Figure 1b draws the experiment as a pipeline: input, model, act, measure, adjudicate. The interpretability bracket covers only the model box. Downstream error is marked at measurement and adjudication. A provenance record that can actually catch those errors has to span the whole chain: prompt, model and tools, plan, reagent information, raw data, and verdict. Raw data has to be kept. A wrong verdict is reopened along that chain.
Figure 1c attacks the idea that a complete log is enough. A sealed log can list every run as PASS, lock the file, and still leave a wrong verdict in place because nothing can be reopened. A re-openable record is also complete, but run 0512 can move from PASS to FAIL once the retained spectra are inspected. Completeness without correctability is a scientific dead end.
Figure 1d states the operating constraint. The schematic scale is about 100 validated-and-recorded runs per day against about 1,000 of execution capacity. Trusted operating rate is capped by validation plus recording. The hatched remainder is unvalidated excess: unaudited risk. An autonomous lab that fills the robot's calendar while the record lags is manufacturing unchecked noise.
The reference list pins the claim to chemical history and ML debate. The Leadbeater, Arvela, and Buchwald palladium-contamination papers show that a published "successful" reaction can be trace metal in the base, not the method. Rudin 2019 is the "don't explain black boxes, use interpretable models" position; McCloskey 2019 showed that attribution on molecular models latches onto spurious features. Read together, interpretability is neither sufficient nor reliably truthful. FAIR and the EU AI Act supply the institutional language for treating provenance as the trust interface. Sutton's Bitter Lesson is cited in the same spirit: do not bet the field on hand-built interpretable structure.
No new experiment, no comparison table. The only numbers are the schematic throughputs in Figure 1d, about 100 recorded-and-validated runs per day versus about 1,000 of execution capacity. Without the body, those two figures should be read as illustration, not as a measured lab rate.
Checkable facts: Comment, pages 804-807, online 20 August 2026; Carnegie Mellon chemical engineering, chemistry, and machine learning; NSF Center for Computer-Assisted Synthesis grant 2202693 and a GSK postdoctoral fellowship.
For anyone wiring an LLM into a wet lab, the acceptance test changes. A log that can only prove "the system said PASS" cannot ground a scientific claim. Interpretability tools still belong in the stack, but they cover the model, not reagent lots, sensor drift, or the adjudication rule.
The engineering bottleneck moves from reactions per day to how many runs can be validated and written into a re-openable record. On the figure's roughly 10x gap, the record subsystem is the capacity limit. Autonomy used to mass-produce uncorrectable conclusions would amplify existing literature bias. It would not correct it.
The abstract's claim that autonomy can become a corrective for the literature's biases, rather than a threat to rigor, has a prerequisite: the record must be re-openable. Miss that, and an agent lab is a faster paper mill.
The paywall hides the four-page prose. The abstract and figure make the thesis clear; they do not show whether the authors specify a field schema or an API for provenance. Do not treat this Comment as a drop-in protocol.
The 100-versus-1,000 scale has no error bar and no cited source. Using it as a capacity formula overreads the figure.
Competing interests are explicit. MacKnight and Gomes co-founded evals, a consultancy for scientific evaluations of frontier models, and are currently affiliated with X, the Moonshot Factory (formerly Google X). A record-first norm lines up with that work. Discount accordingly.
The Comment also leaves a hard question open: who supplies the roughly 100 validated records per day, a person or another model. If the validator is itself a model, the provenance chain just grew another layer that also has to be recorded.