2026-10-09
A GPT-5.2 review of 3,967 open-access TRIPOD-citing prediction papers finds code sharing in 12.2%. Among 380 repos, 37.6% list dependencies and 3.9% include tests.
A clinical prediction model is hard to audit if nobody else can rerun the analysis. Preprocessing, fitting, tuning, and evaluation live in code, and manuscripts rarely capture all of it. TRIPOD, and its 2024 update TRIPOD+AI, tightened what authors must report. Item 18f of TRIPOD+AI encourages sharing analytical code. It does not say what a usable repository contains: a README that states purpose and expected outputs, pinned dependencies, or fixed random seeds.
Earlier measurements were narrow. Only 5.3% of Python notebooks linked from PubMed Central ran end to end. At the Medical Imaging with Deep Learning conference, 22% of public repositories were judged repeatable. About a quarter of 160 deep-learning computational pathology papers released code. Those studies show that a sharing statement and a reusable repository are different things. They do not describe the papers that already cite prediction-model reporting guidelines. This scoping review is stage 1 of TRIPOD-Code, the extension meant to set that standard. It measures the baseline the later Delphi consensus will argue over.
The cohort is every PubMed article citing TRIPOD (6,762) or TRIPOD+AI (411) as of 11 August 2025. After 515 duplicates were removed, 6,658 records remained. Full text had to come from the PMC open-access API, which dropped 1,407 papers. GPT-5.2 (snapshot 2025-12-11) then decided whether each remaining paper developed, validated, or updated a multivariable prediction model: two or more predictors estimating an individual's probability of a current condition or a future outcome. That screen removed 1,284 papers and left 3,967.
The same schema pulled code-availability statements, repository links, and the country of the first author's institution. Two reviewers with prediction-model and clinical-AI experience independently labeled 500 random articles. Against those labels the model reached a weighted F1 of 0.97, and repository-link extraction accuracy of 92.3%. Prompts were frozen after a small pilot.
A retrieval tool accepted GitHub, GitLab, Gitee, Zenodo, Figshare, OSF, and DOIs that resolve to those hosts, always on the default branch. It showed the model the full file tree first, then README and source files in full. Large binaries were dropped. Other files were cut at 3,000 tokens. GPT-5.2 scored 14 reproducibility features fixed in advance by the TRIPOD-Code executive committee: README presence and adequacy, modular structure, dependencies and version pins, license, link to the paper, citation metadata, tests, hardware notes, data or sample inputs, and seed control when a stochastic step was present. Conditional features were scored only when the precondition held.
The same two reviewers labeled 35 repositories. Weighted F1 was 0.83. README, LICENSE, and tests scored F1 of 1.0. Adequacy of documentation scored 0.71.
"Shared code" was defined broadly so the headline rate would not be biased downward. An appendix (75 papers) counted. A link to an unsupported host counted. Only repositories that downloaded and contained at least one non-empty source file entered the quality review. That review covers 380 repositories. Private repos, broken URLs, and profile pages stayed out.
Of 3,967 papers, 482 (12.2%) reported shared code. Each publication year multiplied the odds by 1.12 (95% CI 1.07-1.18). The share moved from 6.3% in 2015 (1/16) and 4.5% in 2016 (2/44) to 14.5% in 2024 (119/820) and 15.8% in 2025 (74/468).
Papers citing only TRIPOD+AI shared code at 29.2%, against 11.4% for papers citing only TRIPOD (OR 3.20, 95% CI 2.23-4.59). Adjusting for year left an OR of 2.61 (95% CI 1.79-3.82). In 2025 alone the split was 29.7% (30/101) versus 11.8% (42/356). The explicit checklist item tracks with more sharing. It does not make sharing the default.
China contributed 32.4% of papers, the United States 11.6%, and the United Kingdom 9.1%. Sharing rates differed by country (Pearson χ²(26) = 134.59). The highest rates were Finland 33.3%, Belgium 26.7%, and Israel 23.1%, among countries with more than 10 papers. Exact denominators for those three are not stated in the text. Journals differed too. Among 71 journals with more than 10 included papers, 17 had a sharing rate of zero. The top rates were Nature Communications 71.4% (n = 14), npj Digital Medicine 57.1% (n = 28), PLOS Digital Health 52.6% (n = 19), and PLOS Medicine 52.6% (n = 19). npj Digital Medicine tells authors that editors may decline a manuscript if important code is unavailable.
Wording was not standardized: 94.4% of phrasings appeared once. Statements sat most often in data availability (232), methods (211), and supplements (84). Most papers had a single statement (68.6%).
Of 447 repository links, 27 used unsupported hosts, 38 did not resolve, and 2 resolved to empty repositories. Of the remaining 380, 83.4% were on GitHub, 3.6% on OSF, 3.4% on Zenodo, and 1.6% on GitLab.
| Feature | Share of 380 repositories |
| README present | 80.5%; README stating purpose and expected outputs, 52.1% of all repos |
| Modular, structured layout | 42.4% |
| Dependencies declared | 37.6%; version-pinned, 21.6% of all repos |
| Data or a sample dataset | 37.1% |
| Link to the paper | 36.8% |
| License | 35.3% |
| Formal citation metadata | 19.5% |
| Hardware requirements | 8.4% |
| Tests | 3.9% |
| Random seed fixed, if stochastic | 64.3% of 286 repos with a stochastic component |
Python appeared in 190 repositories and R in 189. They are not mutually exclusive. The odds a repository contained Python rose 31% per year (OR 1.31, 95% CI 1.16-1.49). The odds it contained R fell (OR 0.85, 95% CI 0.75-0.96), with FDR adjustment across the two tests. Python passed R in 2023 and accounted for more than half of language occurrences in 2025.
No journal led every criterion. Averaged across criteria, PLOS Digital Health sat 11.9% above the global mean. The lowest journal scored zero on seven criteria. A high sharing rate did not imply a strong repository.
TRIPOD-Code will use a Delphi process to mark features as essential, recommended, or context-dependent. This paper is the empirical picture that process starts from. For anyone building clinical prediction models, 12.2% is the relevant base rate inside open-access papers that already cite the reporting guideline: about one in eight shares code. Inside the repos that could be inspected, 35.3% carry a license, 21.6% pin dependency versions, and 3.9% include tests. Code with no license is legally uncertain, so a public URL is not the same as permission to reuse it.
Journal policy moves the sharing rate and does less for repository quality. The four highest-sharing journals all give explicit instructions, yet journals under the same publisher-wide policy still diverge. The practical claim is narrow: requiring deposition does not produce a rerunnable analysis. Asking peer reviewers to audit repositories at scale is unlikely given current review load. Whether the same LLM pipeline should assist that audit is left open.
The per-paper labels and the review code are in the GitHub repository thomas-sounack/TRIPOD-Code, archived as 10.5281/zenodo.22102059.
The authors flag two limits. Misclassification by the LLM means the percentages are trend indicators, not exact census counts. The cohort is restricted to papers that cite TRIPOD or TRIPOD+AI and that PMC open access can retrieve. Those papers are already unusually attentive to reporting guidance, so 12.2% may overestimate sharing in prediction-model research at large. Results should not be read onto STARD or CLAIM.
Several further discounts follow from the design. Repository labels were checked on 35 repos, and the subjective documentation item scored F1 0.71, so "modular" and "adequate README" should not be treated as precise. The 12.2% numerator includes appendices and links the tool never opened, which inflates accessibility. The 1,407 non-open-access papers are simply absent. Country is the first author's institution, with collaboration networks and journal mix left entangled.
Competing interests are disclosed and matter for the journal ranking. Gary Collins and Karel Moons lead TRIPOD and TRIPOD+AI. Leo Anthony Celi is editor in chief of PLOS Digital Health, and Tom Pollard has served on its editorial board. Hyeonhoon Lee is an associate editor of npj Digital Medicine. Those venues are among the highest sharers.
The hardest gap is that nobody reran the repositories. A test file in 3.9% of repos is not a successful reproduction. The notebook rerun study cited in the paper managed 5.3%.