Model hit 97% accuracy, ran in production for a year — it was useless

TajyMany · x · 2026-10-02

An ML engineer shares a horror story: a non-ML engineering team built a classification model that hit 97%+ accuracy, earned promotions, and ran in production for about a year. After a reorg, the ML team analyzed it and found it was no better than random chance. Root causes: useless features, a broken train/test split with heavy overlap, no online performance monitoring, and no review by anyone with real ML experience. Months of development and money wasted — a cautionary tale on ML fundamentals.

Original post →

More from Research

Research channel →