Mathematicians debate whether peer review can survive the flood of AI-generated papers
Fields Medalist Timothy Gowers (Twitter account littmath) and mathematician Obhishek Saha engaged in a heated debate over "where peer review goes when AI produces papers at scale." Gowers revealed he had received submission emails from someone who generated 36,000 papers in one go, arguing that insisting on human-led review at that content scale is "unrealistic"—the traditional mechanisms of author reputation and endorsement by others will break down. A few years from now, if someone endorses or produces 1,000 papers, those signals will mean nothing. The debate drew wide attention for its concrete cases and sharply opposed positions.
Confirmed
- Gowers did receive email submissions from someone who generated 36,000 papers with AI; this is his direct evidence that the review system is unsustainable
- Gowers's core argument: humans cannot review all AI-involved papers, and academia needs some kind of costly signal—such as time invested by the author—to judge what's worth reading before reading it
- He clarified his stance: traces of AI writing in a paper are fine per se, but undisclosed AI use or fully AI-generated papers are negative signals in his view
- Saha's position: judging whether a paper is "entirely AI-generated" is hard or impossible; equating "used AI to write" with "didn't put in enough effort" is a fundamental error—he uses AI at every stage of his own writing workflow
- Saha argues traditional signals (author reputation, endorsement, sloppy writing) still work, and LLMs can help quickly summarize papers and assess mathematical correctness; human reviewers can still lead in the future, just being pickier at acceptance and gating at the topic-selection level
- Commenter kfountou suggested concrete mechanisms: authors submit shareable links to their model interaction logs, or independent agents assess the human contribution in a paper
- Saha also offered a counterintuitive observation: papers scoring 0% on Pangram detection, with no visible AI traces and mediocre prose, are actually a bad sign—the author didn't even spend time polishing with AI
- Another mathematician, Daniel Litt (in follow-up discussion unrelated to Saha), said most predictions from him and his interlocutors will be falsified, but he's certain the math profession will change dramatically and we should stay open to empirical evidence
Unconfirmed
- Details of the 36,000-paper submission come solely from Gowers's account; no original emails or further corroboration have surfaced
- Neither side offered actionable criteria for distinguishing "undisclosed AI use" from legitimate AI-assisted writing in practice
Why it matters
- The debate strikes at the foundations of academic publishing: as content generation cost approaches zero, how do reputation, endorsement, and peer review mechanisms built on scarcity survive—a question every discipline faces
- The disagreement is fundamentally between "introducing new costly filtering signals" versus "keeping old signals and fighting AI with AI"; how it resolves will shape future journal policies, review division of labor, and the institutional design of academic evaluation
2026-10-06 ~ 2026-10-07 · 13 related posts
Primary sources
- Timothy Gowers: we need pre-reading signals to triage AI-generated scientific papers — littmath ·
- littmath reveals emails from someone who generated 36,000 papers in the AI refereeing debate — littmath ·
- The AI refereeing split: humans pick papers, LLMs verify the math, argues @ObhishekSaha — ObhishekSaha ·
- [source] Timothy Gowers: we need pre-reading signals to triage AI-generated scientific papers — littmath · 2026-10-06
- Proposal: shareable chat logs to gauge human contribution in AI-written papers — kfountou · 2026-10-06
- Counterintuitive Take: Papers Refusing Any AI Polish May Be the Bad Sign — ObhishekSaha · 2026-10-06
- littmath clarifies: undisclosed AI use is a negative signal, academia needs a costly signal — littmath · 2026-10-06
- Researchers split on whether AI usage should signal paper quality — ObhishekSaha · 2026-10-06
- littmath: relying on AI-detection to filter papers is de facto the end of human refereeing — littmath · 2026-10-06
- [source] littmath reveals emails from someone who generated 36,000 papers in the AI refereeing debate — littmath · 2026-10-06
- 'People generated 36000 papers': AI spam forces a rethink of peer review — ObhishekSaha · 2026-10-06
- [source] The AI refereeing split: humans pick papers, LLMs verify the math, argues @ObhishekSaha — ObhishekSaha · 2026-10-06
- Mathematician littmath: peer review can't survive a world where one person generates 36,000 papers — littmath · 2026-10-06
- Mathematician argues against AI-usage as a costly signal for paper credibility — ObhishekSaha · 2026-10-06
- Mathematician Daniel Litt: the math profession will change enormously, but nobody knows how — littmath · 2026-10-07
- Mathematician littmath: We Need Pre-Reading Signals to Filter AI-Generated Papers — littmath · 2026-10-07