You Can't Train a Model on a Discovery: Anomaly Detection as the Scientist's Shortlist
bravo_abad · x · 2026-10-07
- Supervised learning can't target discoveries: by definition there are no labeled examples of something nobody has seen. The workable shift is anomaly detection—learn what normal looks like, flag what doesn't fit.
- Isolation Forest does this with almost no assumptions: random trees cut data on random features; points isolated in few cuts are anomalous. Normality is learned implicitly through how hard each point is to separate.
- The scientific catch: anomalous doesn't mean interesting. Over thousands of DFT calculations, top-ranked hits mix metastable structures with unconverged runs, wrong spin initializations, or parsing bugs.
- The division of labor defines the discovery loop: the model ranks what breaks normality, the scientist reviews only what survives. In an autonomous lab, this step decides when to pull a human in.
More from Research
- LiteReality-Agent turns images into simulation-ready 3D scenes for MuJoCo robot planning — elliottszwu · 2026-10-07
- Andrew Davison: robots need object-based SLAM, not scan-then-fit reconstructions — AjdDavison · 2026-10-07
- CtrlCache Speeds Up Interactive Video World Models 1.21–1.41x Without Retraining — Shangye Song · 2026-10-07
- Training-Free Accent Analogy Guidance Boosts Speaker Similarity in Cross-Lingual Voice Cloning — Yoomee Cho · 2026-10-07
- Source Attribution of Synthetic Data Hits 98.7% Accuracy but Falls to 29% After Style Rewriting — Joss Armstrong · 2026-10-07
- Physicist finds fractal patterns (D 1.3-1.5) cut stress response by up to 60% — aakashgupta · 2026-10-07