2026-08-25
At ICML 2023, 1,342 authors ranked 2,592 papers. Over 16 months, top self-ranked papers drew twice the citations of their lowest and beat review scores on highly cited work.
Peer review at AI conferences is being crushed by volume. ICML submissions rose from 1,676 in 2017 to 12,107 in 2025; NeurIPS went from 3,240 to 21,575 in the same window. Experienced reviewers did not scale. Conferences lean on students without prior papers at the venue, and on language models as reviewing aids. In the NeurIPS 2021 consistency experiment, about half of the accepted papers would have been rejected by a second independent committee. Review can still catch factual errors. Under tight deadlines it is easier to score a benchmark bump than to judge whether a paper will still matter in three years.
Buxin Su, Bingxin Zhao, Weijie Su and colleagues at the University of Pennsylvania, with coauthors at Yale, UNC Chapel Hill, NYU and Princeton, tested a different signal: authors who submitted more than one paper to the same conference, ranking those papers against each other. Absolute scores invite inflation. A ranking forces a relative call, and under some game-theoretic conditions that comparison is incentive-compatible. With ICML 2023 organizers' approval, the survey closed before reviews were released.
OpenReview hosted 6,538 submissions from 18,535 authors. The day after the 26 January 2023 deadline, every author received an official OpenRank.cc survey, which closed on 10 February, before reviews went out. 5,634 people responded (30.4%). Of those, 1,342 authors with multiple submissions produced rankings covering 2,592 papers (39.6%). The form stated that rankings would not be shown to coauthors, reviewers, or area chairs, and would not affect 2023 decisions. Authors dragged papers by perceived quality. They were not asked to predict citations or acceptance.
Citations came from Semantic Scholar between 23 July 2023 and 22 November 2024, about 16 months. A record counted as a match only if both title and author-list Levenshtein distances were at most 20. The analysis set is 797 authors and 1,527 unique submissions. Mean citations were 14.73, median 4, 75th percentile 11, maximum 979.
The main comparison is within author and within decision: each author's highest-ranked versus lowest-ranked paper, separately among accepted (poster or oral) and rejected or withdrawn work. 280 authors had multiple accepts; 343 had multiple rejects. The baseline is the same author's pre-rebuttal and post-rebuttal scores on a 1-10 scale. A negative-binomial regression on 2,096 author-submission rows controls for scores and confidence, final decision, portfolio size, number of reviews, and whether the ranking was complete. A 500-split balanced holdout tries to pick the top citation quartile.
Highest-ranked papers averaged 19.99 citations, about twice the lowest-ranked mean (paired Wilcoxon, P < 0.001). Split by decision, the ratio barely moves:
| Group | Self-rank high | Self-rank low | Ratio |
| Accepted | 27.64 | 13.82 | 2.00 |
| Rejected | 12.42 | 6.36 | 1.95 |
Among accepted papers, post-rebuttal scores also separate citations (P = 6.66×10⁻³). Among rejected papers, neither score measure is significant; self-rankings remain so (P < 0.001). When authors ranked at least three papers inside a decision group, accepted means by rank 1, 2, 3, last were 41.57, 19.39, 14.32, 12.83.
The tail is where rankings shine. Of 22 submissions with more than 150 citations, 17 were ranked first by at least one author; 15 were accepted and 7 rejected, including 4 rejects still ranked first. Share above 35 citations: 17.14% of high-ranked versus 6.79% of low-ranked; above 100: 4.29% versus 1.43%. The most-cited accepted high-ranked paper reached 979, nearly three times the low-ranked accepted maximum of 331.
In the fully controlled negative-binomial model, moving from an author's lowest- to highest-ranked paper multiplies expected citations by 1.81. Holdout precision/recall is 0.318/0.307 for self-rank indicators versus 0.307/0.297 for post-rebuttal scores. That is incremental signal, not a standalone classifier, and the paper says so. GitHub stars move the same way: 755.15 versus 347.48 among accepted papers with linked repos. GPT-4o mini ranking 263 accepted pairs prefers papers with 21.69 mean citations versus 19.61; author ranks on the same pairs are 27.00 versus 13.82.
The direction survives dropping self-citations, slicing by narrow score bands, comparing arXiv dates (4 March versus 8 March 2023), and looking at the first two months after the conference (accepted high 1.65 versus low 0.69), before most citing papers could have been written in response to ICML publicity.
For conference chairs this is a nearly free side channel. Authors already know which of their papers they consider stronger; time-pressed reviewers may not. The paper wants rankings as a complement: award shortlists, or a flag when author judgment and scores collide. It does not want them as a replacement for review. ICML 2026 already shipped a version of this. An isotonic mechanism folds rankings into projected scores; about 1,000 high-disagreement submissions were surfaced to area chairs, roughly one per chair, who could order emergency reviews or reread without being told the sign of the discrepancy.
For authors the numbers are colder. Citations are not correctness. Seven papers above 150 citations were rejected. Self-rankings predict whether later work will point at a paper, not whether it should have been accepted. Benchmark deltas remain easier to grade. Papers that authors themselves rank higher often have a longer conceptual bet.
The sample is voluntary and restricted to authors with multiple submissions, so it does not speak for single-paper authors or other fields. Everything already cleared an author's bar for submission; unwritten ideas are out of frame. Observational controls cannot kill residual confounding: an author's "best" paper may also be the one they advertised harder or that sat closer to a hot topic. Preprint dates and early citation windows were checked; that is not exhaustive. If rankings start to affect decisions, incentives change, and 2023's confidential survey is no longer the right model of truth-telling.
The citation window is 16 months, which punishes slow-burn theory. GitHub stars track how software-heavy a subfield is; the authors treat them as descriptive. The GPT-4o mini test is a diagnostic, not an LLM-reviewer bake-off. Follow-up experiments at ICML 2024/2025 and NeurIPS 2025 exist, but those papers have not accrued enough citations yet. The conceptual risk is converting scientific promise into citability and quietly retargeting the conference at that proxy.