S-Agent gets 46.4% on MMSI-Bench by turning spatial reasoning into action chains
机器之心 · wechat · 2026-07-24
South China University of Technology’s S-Lab and Ropedia propose S-Agent, a spatial reasoning framework that turns spatial understanding from one-shot answering into an action chain.
- A VLM reads the question, plans the next step, and delegates geometry-heavy work to specialized models such as DA3.
- The system uses a three-level tool stack: 2D perception, 3D alignment, and spatial experts that convert raw geometry into usable evidence.
- The paper reports 46.4% on MMSI-Bench zero-shot and 60.0% on ViewSpatial-Bench; with S-300K trajectory distillation, an 8B model reaches 41.6 / 46.8 on MMSI/ViewSpatial.
- The authors argue that spatial intelligence should be evaluated as a traceable, reusable agent loop that can accumulate evidence, not as a single answer guessed from a video.
More from Multimodal
- Physics-IQ audit says a third of prompts were ambiguous and 30% of videos inflated scores — DynamicWebPaige · 2026-07-25
- Physics-IQ audit finds ambiguous prompts and artifacts can reshuffle video-model rankings — DynamicWebPaige · 2026-07-25
- A copy-paste prompt for chibi 3D kawaii character generation — cocktailpeanut · 2026-07-25
- AI-generated chibi dolls turn into a collectible-style character set — aziz4ai · 2026-07-25
- AI turns Baki into a live-action style demo — aitrendz_xyz · 2026-07-25
- Microsoft puts MAI-Image-2.5-Flash into Bing Image Creator by default — JordiRib1 · 2026-07-25