2026-08-24
Du et al. argue AI scientists should be scored on discovery episodes that log hypotheses, execution, and interpretation, keep failures, and separate discovery from rediscovery. The abstract reports no new experimental numbers.
Scientific ability is being turned into exams. Benchmarks such as GPQA and Humanity's Last Exam ask whether a model can answer a hard question with a known key. Scores rise. Discoveries do not arrive on the same schedule.
Discovery is a sequence of decisions under uncertainty: make a hypothesis testable, run a computation or experiment reliably, then interpret noisy, failed, or null evidence with restraint. Answering a hard item covers only one slice of that loop. Yuanqi Du at Microsoft Research New England, with coauthors at Stanford, Edison Scientific, the Allen Institute for AI, Deep Principle, and UC Berkeley, argue in a ChemRxiv Perspective that current benchmarks can test whether models answer difficult scientific questions, not whether they can advance science.
This write-up uses the abstract and the bibliographic record only. The ChemRxiv full text was blocked by the site's bot protection. Case studies, figures, and any operational protocol in the body were not checked and are not filled in below.
The proposed unit of evaluation is a discovery episode, not an isolated problem. An episode should record the evolving scientific state, the actions taken, the observations produced, the revisions made, and the provenance needed to reproduce and audit the process.
That record splits three coupled capabilities.
Scoring changes with the unit. Trajectories get scores, not only terminal answers. Failures and null results stay in the record instead of being dropped. Genuine discovery is separated from rediscovery of material already in training data.
The same endpoint can come from recalling a paper, overfitting, or a lucky hit. Trajectory plus provenance is what splits those cases. The evidence of discovery lives in the process, not in the last sentence.
This is a position paper. The abstract reports no new experimental numbers and no head-to-head table against GPQA or other science benchmarks. The deliverable is an evaluation frame, not a leaderboard.
If the body surveys existing scientific-agent benchmarks, defines how to slice episodes, or specifies how to aggregate trajectory scores, that material was not retrieved and is not restated here.
Teams building AI scientist systems or evals can keep adding harder exam items. Exams saturate. Discovery does not follow the same curve. The shift in objective is concrete: log state, action, observation, and revision; let failures enter the score; treat contamination as "was this finding already in the training data," not only "was this question seen."
The smallest usable change is a log, not a new exam. A leaderboard that reports only final-answer accuracy is the thing this paper is arguing against. The author list spans closed-loop lab work and agent evaluation, which is a signal that both sides of that split already feel the mismatch.
It remains a normative piece. There is no released item set to run. The value is a change in what gets scored, not a ranking anyone can ship tomorrow.
The abstract says how AI scientists should be evaluated. It does not say how the scores would be computed. Episode boundaries, trajectory aggregation, and a detector for rediscovery are unspecified. Without those, the frame can stall as a slogan.
The full text was not retrieved, so it is unknown whether the body already operationalizes any of this. A Perspective also supplies no new data for the implied claim that episode-level scoring would change model-development priorities.
Turning discovery into a scorable episode can also flatten open exploration into short, easy-to-log loops. The abstract insists on keeping failures. It does not say how to stop teams from running only the experiments that audit well.