A Beautiful 2D Embedding Is Not Quantitative Evidence, Warns AI-for-Science Researcher
bravo_abad · x · 2026-10-05
An AI-for-Science researcher cautions that t-SNE and UMAP plots — used across crystal structures, molecules, spectra, single-cell profiles and more — should never be treated as quantitative evidence. These methods mainly preserve local neighborhoods and necessarily distort global geometry when squeezing data into 2D: distant-looking clusters in t-SNE aren't necessarily very different, cluster sizes and densities are unreliable, low perplexity can reveal pure noise, and much of UMAP's apparent global structure comes from initialization. Plots shift with perplexity, nneighbors, mindist and the random seed.
The recommended workflow: use 2D maps to explore neighborhoods, subpopulations, outliers and mislabeled entries — then validate in the original space via neighbor computations, cross-seed checks, PCA comparison, and physical ground truth like composition, symmetry, DFT energies and experiments. Principle: generate hypotheses with t-SNE/UMAP, don't turn visual distances into scientific measurements.
More from Research
- Shanghai AI Lab Proposes SCALE: Entropy-Gated Control to Reverse SFT Features — Shanghai-AI-Laboratory · 2026-10-05
- UCVG.cpp: Generate Control Vectors for Any LLM From a Single Prompt Pair — Egor4more · 2026-10-05
- How many digital minds on one GPU cluster? Synthese paper probes interwoven AI consciousness — burny_tech · 2026-10-05
- New preprint argues 'adaptive reframing'—revising problem representations—is a key unstudied dimension of intelligence — ValerioCapraro · 2026-10-05
- Stop Reading the Hessian as a Matrix: Eigenvalues Are Local Curvature of the Loss Surface — techNmak · 2026-10-05
- ML Conference AC warns: papers that are 'incomprehensible' will be desk rejected — tyrell_turing · 2026-10-05