AI for Science Reproduction Fails: New Data Makes Battery Electrolyte Model Worse
bravo_abad · x · 2026-07-31
Reusability reports in ML for science are severely undervalued. An independent team rebuilt a published predictive-generative model for liquid electrolyte formulation and stress-tested it until it broke.
While reproduction was clean and random seeds barely affected results, they uncovered critical failure modes. In pretrain-then-finetune pipelines, data importance is highly asymmetric: cutting the pretraining set to 10% barely impacted conductivity prediction, but cutting the fine-tuning set to 10% caused a significant performance drop. Counterintuitively, adding data points from a new regime made the model worse. These findings transfer to almost any pretrain-then-finetune pipeline.
More from Research
- AdaMAST: Boosting AI Agent Performance with Failure Taxonomies — abeirami · 2026-07-31
- Hardcore Systems Engineering in Kimi K3 Paper: Compilers and Chip Design — DynamicWebPaige · 2026-07-31
- RBC Borealis Releases Comprehensive Tutorial Series on ML Math — SimonPrinceAI · 2026-07-31
- Circulation Journal Reviews Generative and Agentic AI in Drug Discovery — james_y_zou · 2026-07-31
- Qwen Releases Technical Report for Qwen-Audio-3.0-Gen-Preview — udmrzn · 2026-07-31
- Demystifying the Core Math Behind Large Language Models — udmrzn · 2026-07-31