Six-month arXiv audit: substantive AI use in math papers jumped from 1.4% to 14%

The Gold Rush in AI4Math: Where Are We Now?

Jiashun Jin, Zheng Tracy Ke, Bingcheng Sui

stat.AP, cs.CY, cs.DL, math.HO

2026-08-25

Audit of 32,944 math arXiv papers: substantive AI use rose from 1.39% in March to 14.09% by mid-August; 71% of 717 named open problems claimed fully resolved.

What problem this solves

Capability demos are no longer scarce. AlphaGeometry, AlphaProof, the First Proof project, and a handful of 2026 public claims (a Hamiltonian-cycle construction, a counterexample to Erdős's unit-distance conjecture, Astra's ten-problem report) all ask whether AI can do research math. A different question has almost no systematic data behind it: how often mathematicians are already using these systems, in which fields, and for what.

This paper audits every arXiv submission from 1 March to 20 August 2026 whose primary or secondary category includes Mathematics: 32,944 PDFs. The unit of observation is an explicit author disclosure in the manuscript. Undisclosed use is invisible, so every prevalence number is a lower bound.

Method

All 32,944 PDFs were parsed to text. Files longer than 300,000 characters kept only the first and last 150,000; 1,007 records were truncated that way. GPT-5.6 Terra, low reasoning effort, 100 parallel workers, then filled a fixed questionnaire. The questionnaire has three blocks:

Every judgment had to quote the manuscript and attach a confidence score from 1 to 3. Human validation is still running: a keyword screen produced about 5,000 candidates, then a volunteer team of more than ten people has been filling the same form on a sample. The numbers below are an LLM reading disclosures, not mathematicians checking proofs.

Results

3,575 submissions disclose some AI use; 1,712 of those have at least one substantive label. Confirmed disclosure rose from 4.75% of math submissions in March to 24.14% through 20 August. Substantive use rose from 1.39% to 14.09%.

CategoryPapers
Proof construction1,225
Formalization and verification427
Other research assistance289
Problem formulation146
Language and formatting2,108
Code and computation1,229

Language editing is still the single most common bin, but proof construction already exceeds a thousand papers. Combinatorics leads in count with 397 substantive papers. Metric Geometry leads in rate: 40 of 305 submissions, 13.11%. Number Theory sits at 135 papers and 7.95%. Large fields such as Numerical Analysis and Analysis of PDEs have much lower substantive rates, 1.91% and 2.21%.

Among 717 named open-problem records tied to substantive use, 510 are labeled fully resolved (329 proved, 181 disproved), 103 still open, 93 partial, 11 unclear. That is a 71% fully-resolved share. The labels come from the authors' own wording. They are not independent verification. The fifty earliest proposed-year records that the pipeline marks as fully resolved run from the 1913 Ramanujan–Nagell theorem to Talagrand's convolution conjecture in the late 1980s; 14 of those fifty are Erdős-related.

People, institutions, countries, and vendors are concentrated. The 1,712 substantive papers involve 2,764 canonical authors and 744 university-level institutions. 83.6% of authors appear once. Under fractional author weights, the United States accounts for 33.7% of recognized country weight and China 32.9%, about two-thirds together. Manuscript counts by institution: MIT 35, USTC 34, Stanford 33, Peking University 31. Author-level Gini is 0.449; institution-level Gini is 0.606. OpenAI is named in 2,148 of 3,575 confirmed-use papers (60.08%), Anthropic in 776 (21.71%), Google in 334 (9.34%), DeepSeek in 80 (2.24%). A paper can name several vendors, so the shares do not add to 100%.

Why it matters

The debate about whether AI will change mathematics now has a reproducible baseline, not only a pile of demos. In six months, disclosed substantive use moved from a rounding error to more than one paper in seven. The fields that moved first (combinatorics, number theory, metric geometry) are the same ones where automated reasoning has been landing results. Watching contest scores or a few famous counterexamples understates the disclosed participation already on the archive.

The same table shows how narrow the early adopters are. OpenAI is named in three-fifths of disclosing papers; the US and China dominate country weight; a short list of universities appears again and again. "The mathematical community is using AI" currently means a small set of authors and labs, and only the part they chose to write into the PDF.

Limitations

The paper lists four caveats of its own. The study measures disclosed use, so 14% is a floor. An LLM classified authors' wording and will mislabel some papers; the human audit is unfinished. Institution and country assignments mix normalization and inference, good for a distribution, bad as a ranking. Open-problem "resolutions" are not peer review, priority research, or a correctness check.

A few more gaps sit in the design. Disclosure intensity is not comparable across papers: polishing prose with ChatGPT and putting a model in the main line of a proof can look similar in a checkbox. The March-to-August spike mixes new models, shifting community norms, and arXiv's disclosure policy; the observational series cannot separate those. A 71% fully-resolved rate among 717 named problems is high enough to discount by default. Duplicate filings, restated old results, and proofs with holes all inflate "resolved." Table 3 is a sample of what AI-assisted math is claiming, not a list of what AI has settled.

Terms

Source

What people are saying

All paper explainers