Manual checks of 105 papers find most reproducibility failures come from missing code and mismatched results
ChenhaoTan · x · 2026-07-23
Manual checks across 105 papers find reproducibility failures are usually mundane, not mysterious
The author says their analysis suggests that the main reasons papers fail to reproduce are often not model capability issues. Instead, the problems are usually practical and procedural.
They give one concrete example: a paper claims its method trains only 0.77% of the base model’s parameters, but the checkpoint actually released trains 6.31%—about 8× more. The smaller figure only works if you assume an adapter one-eighth the size of the one that was shipped.
The attached chart summarizes the most common causes of non-reproduction in their sample of 105 papers:
- Code doesn’t run as shipped: 58
- Numbers don’t match the paper: 48
- Data not available: 42
- No runnable code released: 38
- Trained models not released: 18
- Code doesn’t match the paper: 15
- Relies on legacy or unavailable models: 4
The takeaway: 101 of 105 papers had at least one of these issues, so reproducibility is often blocked by missing artifacts, mismatched code, and inconsistent reporting rather than by raw model capability.
More from Research
- AI could compress decades of biomedical research into days, says Derya Unutmaz — DeryaTR_ · 2026-07-23
- OpenAI and Apollo show models may optimize graders, not user intent — rohanpaul_ai · 2026-07-23
- AI Autonomously Disproves Decades-Old Math Conjectures: The Singularity's Opening Phase — imjustnewatai · 2026-07-23
- Applied Math Dominates AI, But Why Does Gradient Descent Actually Work? — fkasummer · 2026-07-23
- Cursor’s Composer 2.5 looks much worse at reasoning than its Kimi base model — gleech · 2026-07-23
- Lanyon says its neurosymbolic solver is 20–250x faster than frontier models — burny_tech · 2026-07-23