Preprint: stop using PCA on LLM item embeddings — EGA recovers structure 97.6-100% of the time
GolinoHudson · x · 2026-09-30
A new preprint by Garrido, Russell-Lasalandra, Rodríguez-Montoya and Golino delivers a blunt recommendation: stop using PCA to estimate dimensional structure from LLM item embeddings.
The team tested PCA vs. Exploratory Graph Analysis (EGA) across 3 LLMs (GPT-4o, GPT-5.4, Claude Sonnet 4.6), 2 embedding models (OpenAI text-embedding-3-small, Jina v3), 6 known personality dimensions, hundreds of Monte Carlo replications, plus an empirical replication on the Multidimensional Schizotypy Scale.
Key findings:
- EGA recovered the correct 6-dimensional structure in 97.6-100% of conditions; PCA + Kaiser and PCA + Parallel Analysis achieved essentially 0% exact recovery
- For a true 6-factor structure, PCA frequently estimated 8-24 dimensions. The cause: few items create a long tail of "nuisance" dimensions, so eigenvalue-based methods massively overextract
- In their AI-GENIE package, UVA removes redundant items while bootEGA flags items with unstable dimensional placement — redundancy ≠ structural instability, and different LLMs showed different failure profiles
- On the real psychological instrument, EGA came substantially closer to the human-response benchmark than PCA under both embedding models
Related event: Preprint: EGA beats PCA for estimating dimensionality in LLM embeddings(2 posts)→
More from Research
- Diffusion Models Tutorial Accepted to NeurIPS 2026 Alongside 7 Paper Acceptances — mittu1204 · 2026-09-30
- AMB3R-SLAM: Kilometer-Scale Real-Time SLAM on One Consumer GPU, Cutting ATE by 70% — rsasaki0109 · 2026-09-30
- Prefix-Reuse FLOPs: new metric exposes hidden cost of arbitrary context edits in LLM serving — RulinShao · 2026-09-30
- Tencent Hunyuan releases ExplorationBench to measure how AI systems explore — TencentHunyuan · 2026-09-30
- Fully open MolmoAct 2 tops independent robotics benchmark LIBERO-MAX on dynamic robustness — DJiafei · 2026-09-30
- CompVis improves Distributional Diffusion Models: 4.48 FID at 4 steps on ImageNet — CompVis · 2026-09-30