2026-09-02
At ICML 2023, 1,342 authors privately ranked their papers. Top-ranked ones drew twice the 16-month citations of the lowest; review scores failed on rejects.
This Nature Computational Science piece is a News & Views by Bedoor AlShebli (NYU Abu Dhabi) on an ICML 2023 field experiment by Buxin Su, Bingxin Zhao, Weijie Su and colleagues at Penn and elsewhere. The experiment asked authors to privately rank their own multiple submissions, then tested whether that ranking forecasts later citations better than peer-review scores.
Top-venue reviewing is already overloaded. ICML submissions rose from 1,676 in 2017 to 12,107 in 2025; NeurIPS went from 3,240 to 21,575 in the same window. Experienced reviewers have not scaled with that growth, so conferences lean more on graduate and undergraduate reviewers and have started using language models for feedback. In the NeurIPS 2021 consistency experiment, about half of the accepted papers would have been rejected by a second independent committee.
Reviewers on a clock tend to reward a clean benchmark bump. They are worse at judging whether a paper will still matter in three years. Authors know the conceptual depth and longer bet of their own work, but an absolute self-score invites inflation. The test is narrower: if you only ask an author to order their own papers, does that relative judgment carry real information?
ICML 2023 had 6,538 submissions from 18,535 authors. Right after the 26 January 2023 deadline, organizers sent an official OpenRank.cc survey that closed on 10 February, before reviews went back to authors. Authors with multiple submissions dragged their papers into a perceived-quality order. The form stated that rankings would not be shown to coauthors, reviewers, area chairs or program chairs, and would not enter 2023 decisions. Authors were not asked to predict acceptance or citations.
Of 5,634 respondents (30.4% of authors), 1,342 with multiple submissions ranked 2,592 papers (39.6% of the conference). Reviews used a 1–10 scale, recorded both before and after rebuttal. Citations were taken from Semantic Scholar for 23 July 2023 through 22 November 2024 (16 months), matching titles and author lists with Levenshtein distance at most 20. The analysis set is 797 authors and 1,527 submissions with valid citation records. The main contrast is within an author's own portfolio: highest-ranked versus lowest-ranked paper, split by accept versus reject, tested with two-sided paired Wilcoxon signed-rank tests.
The design uses comparison, not scoring. An author cannot mark every paper as best. Weijie Su's earlier isotonic mechanism gives game-theoretic conditions under which this kind of ranking is incentive-compatible.
Across the full sample, top-ranked papers averaged 19.99 citations, about twice the low-ranked mean (P < 0.001). Conditioning on the decision barely changes the ratio:
| Split | High self-rank mean | Low self-rank mean | Ratio | Wilcoxon P |
| Accepted (n=280 pairs) | 27.64 | 13.82 | 2.00 | 6.00×10⁻³ |
| Rejected or withdrawn (n=343 pairs) | 12.42 | 6.36 | 1.95 | < 0.001 |
Review scores do not match that pattern. Among accepted papers, the post-rebuttal high-versus-low gap is significant (P = 6.66×10⁻³); the pre-rebuttal gap is not (P = 0.112). Among rejects, neither score is significant (P = 0.890 and 0.556). Self-ranking is the only signal that holds on both sides of the decision.
The tail is sharper. Of 22 papers with more than 150 citations, 17 (77.27%) were ranked first by at least one author. Those 22 papers are about 1.44% of the analysis set, close to a typical oral rate. At 35 citations, 17.14% of high-ranked papers clear the bar versus 6.79% of low-ranked papers; at 100 citations the split is 4.29% versus 1.43%. The most-cited paper reached 979 citations. Its authors ranked it first; its post-rebuttal score was 5; it was accepted as a poster.
In a fully controlled negative-binomial model (review scores and confidence, decision, portfolio size), moving from an author's lowest to highest paper is still associated with 1.81 times the expected citation count. Over 500 holdout splits, self-rank indicators reached precision 0.318 and recall 0.307 for the predicted top quartile, versus 0.307 and 0.297 for post-rebuttal scores. The lift is small and larger among rejects. GitHub stars move the same way: 755.15 versus 347.48 among accepted papers with linked repos, 89.77 versus 33.27 among rejects. The gap survives dropping self-citations and is not explained by posting time (mean arXiv dates 4 March versus 8 March 2023).
A diagnostic splits the source of the judgment. GPT-4o mini ranked 263 accepted within-author pairs from text alone: preferred papers averaged 21.69 citations versus 19.61 for the rest. On the same pairs, author rankings were 27.00 versus 13.82. The comparative format is not enough; who ranks matters.
For conference organizers this is almost free side information: authors already submitted multiple papers, and dragging an order takes a minute. ICML 2026 has already used it. After pre-rebuttal scores, an isotonic mechanism produced projected scores consistent with author rankings. About 1,000 submissions with a large gap between projected and raw scores were flagged as disagreement for area chairs and senior area chairs, without revealing the direction, so they could request emergency reviews or reread existing ones.
For authors the concrete case is the 979-citation poster with a score of 5. A low score plus a first-place self-rank is not automatically a weak paper. This is an incremental add-on, not a replacement for peer review. The original study is explicit: rankings are targeted side information for award shortlists or extra scrutiny, not a substitute for checking correctness.
Citations proxy influence, not quality. A citation can be critical or fashionable, and 16 months is short for slow-burn work. The sample covers only volunteers with multiple submissions; single-paper authors are absent. A 30.4% response rate likely selects people willing to state a relative preference. Every paper had already cleared the author's own submission bar, so the result does not speak to ideas that were never sent.
The design is observational and cannot kill residual confounding. If rankings start to affect decisions, author incentives change and the incentive-compatibility argument may not hold. The GPT check used GPT-4o mini on 263 accepted pairs; it does not travel to stronger models or the full pool. Follow-on experiments at ICML 2024 and 2025 and NeurIPS 2025 still need time to accrue citations.
AlShebli's News & Views itself sits behind a paywall; only the abstract, a figure caption and seven references were recoverable. The numbers and methods above come from the open-access Brief Communication it comments on (Su et al., 2026).