Researchers find S-Space: multimodal models encode left/right/up/down/depth in a linear subspace
jiqizhixin · x · 2026-09-18
MirroS researchers report discovering S-Space (Spatial Workspace) inside multiple multimodal models: by analyzing intermediate activations, they found objects' left-right, up-down, and near-far positions are written into a stable low-dimensional representation space — a continuous linear subspace spanned by three directions, where reading along each axis yields the object's continuous coordinate. Striking examples include the model inferring an off-screen ball's position from a diving goalkeeper, and the abstract concept "socialist" landing on the left side of the internal space, mirroring the "left wing" spatial metaphor in language.
More from Research
- XGEN unveils generative world simulation prototype; JING tops WBench leaderboard — SucceededMind · 2026-09-18
- Is the Hype Around Type Safe AI's Jev and RLCD Justified? A Skeptic's Take — Haghiri75 · 2026-09-18
- New academic spam: single-authored papers cold-pitching ARR service contributions — anmarasovic · 2026-09-18
- A 'Life Diary' Eval Could Be the Toughest Test Yet for Continual Learning in LLMs — JohnnyNi13 · 2026-09-18
- LLMs got good at text and stayed bad at tables — and it's not just a training-data problem — FamiliarSlide7685 · 2026-09-18
- IFM releases K2-Horizon-7B, a diffusion-augmented LLM claiming lossless 5,200 tokens/s — Zulfiqaar · 2026-09-18