AI for Science Reproduction Fails: New Data Makes Battery Electrolyte Model Worse

bravo_abad · x · 2026-07-31

Reusability reports in ML for science are severely undervalued. An independent team rebuilt a published predictive-generative model for liquid electrolyte formulation and stress-tested it until it broke.

While reproduction was clean and random seeds barely affected results, they uncovered critical failure modes. In pretrain-then-finetune pipelines, data importance is highly asymmetric: cutting the pretraining set to 10% barely impacted conductivity prediction, but cutting the fine-tuning set to 10% caused a significant performance drop. Counterintuitively, adding data points from a new regime made the model worse. These findings transfer to almost any pretrain-then-finetune pipeline.

Original post →

More from Research

Research channel →