S-Agent gets 46.4% on MMSI-Bench by turning spatial reasoning into action chains
机器之心 · wechat · 2026-07-24
South China University of Technology’s S-Lab and Ropedia propose S-Agent, a spatial reasoning framework that turns spatial understanding from one-shot answering into an action chain.
- A VLM reads the question, plans the next step, and delegates geometry-heavy work to specialized models such as DA3.
- The system uses a three-level tool stack: 2D perception, 3D alignment, and spatial experts that convert raw geometry into usable evidence.
- The paper reports 46.4% on MMSI-Bench zero-shot and 60.0% on ViewSpatial-Bench; with S-300K trajectory distillation, an 8B model reaches 41.6 / 46.8 on MMSI/ViewSpatial.
- The authors argue that spatial intelligence should be evaluated as a traceable, reusable agent loop that can accumulate evidence, not as a single answer guessed from a video.
More from Multimodal
- Dev builds interactive 3D product experience with GPT-6 Astra + Hyper3D Rodin — nikola_mr64990 · 2026-09-11
- Using a finisher move on one mosquito with MiniMax H3 MAX — the bug survives — Hailuo_AI · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Lumara AI Film Festival Comes to NYC Oct 26, Top AI Filmmakers to Compete — 0xAllen_ · 2026-09-11
- Pterodactyl Detective: An AI-Generated Proof-of-Concept Trailer — PterodactylDetective · 2026-09-11
- Imperium Game Trailer Showcases AI Video Generation — keaslenyt · 2026-09-11