Manual checks of 105 papers find most reproducibility failures come from missing code and mismatched results

ChenhaoTan · x · 2026-07-23

Manual checks across 105 papers find reproducibility failures are usually mundane, not mysterious

The author says their analysis suggests that the main reasons papers fail to reproduce are often not model capability issues. Instead, the problems are usually practical and procedural.

They give one concrete example: a paper claims its method trains only 0.77% of the base model’s parameters, but the checkpoint actually released trains 6.31%—about 8× more. The smaller figure only works if you assume an adapter one-eighth the size of the one that was shipped.

The attached chart summarizes the most common causes of non-reproduction in their sample of 105 papers:

The takeaway: 101 of 105 papers had at least one of these issues, so reproducibility is often blocked by missing artifacts, mismatched code, and inconsistent reporting rather than by raw model capability.

Original post →

More from Research

Research channel →