A data scientist's two R-squared mistakes that hurt his regression models for 2 years
mdancho84 · x · 2026-09-01
The author reviews two mistakes from his early regression and forecasting work:
- Mistake 1: in-sample metrics. Chasing a high R² on training data led to overfitting and erratic forecasts; switching to cross-validation fixed much of it.
- Mistake 2: maximizing variance explained only. R² measures explained variance, but forecasting wants stability around the mean of unseen data, which calls for different metrics.
After fixing both, his models were roughly 50% more accurate than his peers'.
More from Research
- Stop building remote-controlled robots; sim2real is the key to success — _Stocko_ · 2026-09-02
- CS Professor shares the evolution of their paper-reading stack over the years — CSProfKGD · 2026-09-02
- OpenBind preprint released with new virtual screening benchmark — MoAlQuraishi · 2026-09-02
- Anima Anandkumar launches acceleratedU: neural operators over transformers for physical prediction — AnimaAnandkumar · 2026-09-02
- New blog 'Proofs and Prompts' explores how AI is reshaping mathematics — giannis_daras · 2026-09-02
- EMNLP 2026 Paper: SCALE uses structured case law for legal reasoning — ponguru · 2026-09-02