Preprint: PCA unreliable for LLM embeddings, EGA wins across 3 LLMs and hundreds of simulations
GolinoHudson · x · 2026-09-30
A new preprint argues against using PCA to estimate dimensional structure from LLM item embeddings.
The group benchmarked PCA against Exploratory Graph Analysis (EGA) across:
- 3 LLMs: GPT-4o, GPT-5.4, and Claude Sonnet 4.6
- 2 embedding models: OpenAI text-embedding-3-small and Jina v3
- 6 known personality dimensions
- Hundreds of Monte Carlo replications, plus an empirical replication using Multidimensional Scaling
Bottom line: EGA recovers known dimensional structure more reliably, and the authors recommend dropping PCA for this task.
Related event: Preprint: EGA beats PCA for estimating dimensionality in LLM embeddings(2 posts)→
More from Research
- Missing API for general real-time LLM agents: AsyncLLM preprint sparks interface debate — phill1992 · 2026-09-30
- Cohere Labs to host IOL-AI 2026 wrap-up on why linguistic reasoning still stumps models — Cohere_Labs · 2026-09-30
- UK's Zenithon raises $10M to build world models for extreme physics: rockets, fusion and fabs — roydanroy · 2026-09-30
- Tenstorrent opens bio-model training: OpenFold3 on Blackhole Galaxy nears DGX H200 at quarter the cost — MoAlQuraishi · 2026-09-30
- MIT's Ataraxo AI beats top Stratego players with self-play and decision-time planning — nordicinst · 2026-09-30
- New Research: AI as Tutor Beats AI as Substitute — and No AI — CackleRooster · 2026-09-30