2026-09-05
Gemini-2.0 Co-Scientist lifts hypothesis Elo with test-time compute, hits AML IC50s to 2 nM (KIRA6 10 nM on KG-1a), and recaps an unpublished cf-PICI mechanism in two days.
Scientists are squeezed from both sides. Subfields keep getting deeper, while the jumps that actually move a field often come from stitching across disciplines. The literature is already too large for any person to finish. Tools in the "deep research" genre can summarize what is already known. They stop at the summary. They do not propose a new hypothesis you can take to the bench.
Co-Scientist aims at the next step. A scientist writes a research goal, preferences, and experimental constraints in ordinary language. The system keeps generating, debating, and evolving hypotheses, and it returns experimental protocols. The backbone is Gemini. The architecture is framed as a general scientific engine; the validations in this paper are biomedical.
All agents run on Gemini 2.0. The authors describe the framework as model-agnostic. A Supervisor parses the goal into a research-plan configuration and dispatches work onto an asynchronous task queue. Six specialist agents map onto pieces of the scientific method:
Scientists can inject constraints, seed their own ideas, or kill a direction in natural language at any time. Tools include web search, a private corpus, and specialized models such as AlphaFold. Default gates are goal alignment, plausibility, novelty, testability, and safety.
Across 203 research goals, mostly biomedical with some math and physics mixed in, hypotheses were sliced into ten time buckets. Both best-Elo and top-10 average Elo climb with test-time compute. On the figure the best score moves from roughly 1300 toward 1600, with no saturation in the plotted window. On 15 hard goals written by seven biomedical experts, Co-Scientist eventually outranks contemporaneous Gemini 2.0 Pro Experimental, Gemini 2.0 Flash Thinking, OpenAI o1, o3-mini-high, and DeepSeek R1 on Elo. o3-mini-high and R1 were competitive at much lower compute. Elo here is a self-score, not a wet-lab gold standard.
In a blinded review of 11 of those goals, the same experts ranked Co-Scientist 2.36 on average (1 is best, 4 is worst), with novelty 3.64 and impact 3.09 on a five-point scale. Gemini 2.0 Flash Thinking sat at rank 2.73 with novelty and impact both 3.09; o1 scored 3.55 novelty and 2.82 impact. These are preference ratings on a small sample.
The weight-bearing results are in vitro.
AML repurposing. Experts picked five candidates that already had some preclinical rationale. Binimetinib, pacritinib, and cerivastatin inhibited viability. Binimetinib reached IC50 values as low as 2 nM on most AML lines, much higher on the non-AML lymphoblastoid line TK6. On MOLM-13 the plotted IC50 is 0.01 μM (about 10 nM) for binimetinib, 0.73 μM for pacritinib, and 6.73 μM for cerivastatin. Of three candidates with no prior AML preclinical evidence, the IRE1α inhibitor KIRA6 worked: 10 nM IC50 on KG-1a versus 180 nM on TK6, an 18-fold window; 144 nM on NOMO-1; 1750 nM on MOLM-13 and 870 nM on HL-60. Nanvuranlat and leflunomide barely moved MOLM-13. Seven doublet or triplet combinations were mostly synergistic on MOLM-13 (JNJ-64619178 + selinexor; JQ1 + olaparib + MSA2) and mixed synergistic/antagonistic on TP53-mutant KG-1a.
Liver fibrosis. The system proposed three epigenetic targets. Experts chose matching drugs; two showed anti-fibrotic activity in human hepatic organoids without obvious cytotoxicity. One of those drugs is the already-approved anticancer agent vorinostat.
Antimicrobial resistance. Given minimal background, the system independently proposed in two days that capsid-forming phage-inducible chromosomal islands (cf-PICIs) expand host range by interacting with diverse phage tails. That matched a then-unpublished experimental finding from a collaborating group, a line of work that had taken humans about a decade (2015–2024).
Few multi-agent hypothesis engines have been pushed all the way to wet-lab checks. For anyone doing drug screens, the paper shows a short loop of literature reasoning, expert triage, then a cell-line sanity check, without laying down a full combinatorial wet-lab matrix first. The stack is not glued to one backbone, so a later Gemini generation can be swapped in.
This is not an autonomous scientist. In-vitro activity still sits far from in-vivo or clinical success, across pharmacokinetics, microenvironment, and patient stratification. Elo is self-assigned. The expert ranking covers 11 items. Treat it as a hypothesis accelerator, not a conclusion machine.
The authors are explicit. Knowledge is bounded by open-access literature, so paywalled work and negative results are missing, and irreproducible published claims can be recycled. The backbone still hallucinates. All three wet-lab tracks are preliminary. There is a real risk that model taste homogenizes research directions, and that unreviewed mass output makes the reproducibility crisis worse.
A few further weak spots. Elo self-scores and "the hypotheses got better" are produced by the same debate machinery, which has a self-grading flavor. Expert preference tracks Elo, but n=11, and the paper says that is too small for firm claims. The AML assays measure viability, not mechanism; KIRA6's IC50 spans two orders of magnitude across subtypes, which looks like a biomarker problem rather than a broad-spectrum hit. Key experimental detail for fibrosis and AMR lives in companion Cell and Advanced Science papers; the Nature text retells those results at narrative resolution.