40,000 LLM Agents Simulate a Decade of Academia: Resubmission Alone Triples Reviewer Burden

Science Utopia? Closed-Loop LLM Simulation of Academic Research Ecosystems

Yiqiao Jin, Yiyang Wang, Lucheng Fu, Bing He, Siheng Xiong, Yijia Xiao, B. Aditya Prakash, Josiah Hester, Srijan Kumar, James Evans, Jindong Wang

cs.CL, cs.AI, cs.CY

2026-10-01

SUTO simulates academia with 40,000 LLM agents: resubmission lifts reviewer load from 2.07 to 6.64 papers; doubling output per project adds 63% papers but cuts 10-year survival to 24% from 39%.

What problem this solves

AI is moving into every stage of research: picking topics, running experiments, writing manuscripts, reviewing them. Gains at each stage do not automatically add up to a healthier research system. Acceptance rates, funding rules, and resubmission policies feed back into one another, and the consequences take a decade to surface, yet real academia allows no controlled experiments; nobody gets to randomly change a conference's acceptance rate and watch what happens. Existing work splits into two camps with complementary gaps: AI-scientist style systems handle single tasks with no persistent state, while classic agent-based models of science capture institutions but run on toy rules with no real scientific content. SUTO fills that gap.

Method

SUTO, from Georgia Tech, UChicago, William & Mary and collaborators (code released), runs each simulated year as a six-phase cycle: direction choice and collaboration, submission and citation, peer review, decisions and memory update, funding and attrition, then resubmission.

Scale: 61 worlds, 40,000+ researcher agents across 8,000 institutions, roughly 400,000 publication decisions and 1.2 million LLM-generated reviews.

Results

SettingMetricResult
Population growth only / baselineReviews per active reviewer2.25 / 2.07
Resubmission onlyReviews per active reviewer6.64
Fixed 30% vs ACL-style declining acceptance (28% to 20.3%)10-year submission growth / review workload282% vs 358%; stricter regime adds 14% workload
Submission cap 1 to 2, grants scale with demandAccepted papers / year-10 survival+69%, 49% to 29%
Submission cap 1 to 2, fixed grantsAccepted papers / year-10 survival+63%, 39% to 24%
Explorer / cautious explorer / exploiterAcceptance rate29.8% / 37.2% / 39.2%
Double-blind vs revealed first-author name and institutionReview score (1-5 scale)+0.055 to +0.077

Resubmission is the main engine of reviewer burden. Population growth enlarges the reviewer pool along with submissions, so per-reviewer load barely moves; enabling resubmission alone pushes it to 6.64. A stricter acceptance regime raises total review workload by 14% while first submissions barely change, because rejected papers pile review labor into later rounds. Under growth plus resubmission, resubmissions account for 61% of year-ten submissions.

Publication growth can hide a collapse in participation. Doubling the submission cap yields 63 to 69% more accepted papers, yet year-ten survival falls from 39% to 24% under fixed grants, and from 49% to 29% even when grant capacity scales with demand. In mixed university-industry worlds, industry researchers drop from 87% to 56%. Switching from scaling to fixed grants raises the never-funded share from 22% to 37%, while the Gini among funded researchers edges only from 0.36 to 0.39: scarcity pushes more people out rather than enriching a few.

Cautious exploration is the best-value individual strategy. Cautious explorers, who open new directions near existing expertise, sit 2 points below pure exploiters in acceptance, earn stronger citation impact, and match their roughly 53% year-ten survival; distant explorers fall to 41.1% and have the lowest probability of a high-impact paper (0.62, against 0.81 and 0.84 for the other two strategies). At the ecosystem level, cautious-heavy worlds hold 0.11 higher topic entropy and 8.5 more active directions than exploiter-heavy ones. Two counterintuitive results: papers farther from an author's recent work are much harder to publish (acceptance 48.0% to 21.8% across departure bins), yet the accepted ones see mean citations by age three climb from 0.95 to 1.98; and changing your own direction is not the same as changing the field's, with explorers showing the lowest share of disruption-positive papers (39.8% by CD index).

In untouched large worlds, funding Gini climbs from 0.04 to 0.36 over a decade while the active population shrinks from 5,000 to about 2,530, yet all 53 research directions retain active researchers: resource concentration and intellectual diversity decouple. Narrowly winning early funding shows no detectable career advantage over the following three years. On review bias, revealing the first author's name and institution raises LLM review scores by 0.055 to 0.077, with almost no added effect from disclosing collaborator ties.

Why it matters

This is a wind tunnel for science policy: change one rule, run ten years, observe the long-run consequences. AI-for-science teams can stress-test their tools here, for example whether review infrastructure breaks before productivity gains arrive. Policy readers should take the 'more output, fewer people' result seriously: before cheering AI-accelerated research, ask whether review and funding capacity keeps up. The honest caveat is that every number emerges from LLM agent behavior; this is a thought experiment, not a measurement of real academia.

Limitations

A footnote concedes that the proposal-to-grant pipeline is never simulated; the team assumes publication records correlate with proposal evaluations, which simplifies every funding conclusion. Budget is an abstract resource the authors liken to AI tokens, attrition is sensitive to that parameterization, and the main text offers no sensitivity analysis. All reviews are LLM-generated, so the author-identity bonus measures LLM reviewer bias, not necessarily human bias. There is no systematic calibration against real academic data, and negative results such as the absence of cumulative advantage from narrow early wins depend on the specific funding model and could flip under a different one.

Terms

Source

Related papers

All paper explainers