S-Agent by NTU & Tsinghua: Eliciting 3D Spatial Reasoning in Video via Tool-Use
jiqizhixin · x · 2026-08-05
Researchers from Nanyang Technological University (NTU) and Tsinghua University introduced S-Agent, a spatial tool-use agent designed to elicit reasoning for spatial intelligence.
Core Mechanism:
S-Agent accumulates evidence across video frames, lifting 2D entities into a 3D scene frame to build an egocentric frame of reference and resolve directions.
Performance:
- In zero-shot settings, S-Agent boosts Gemini 3 Pro performance by 1.2% and ranks first on the MMSI-Bench and ViewSpatial-Bench.
- Through Supervised Fine-Tuning (SFT), the distilled S-Agent-8B improves Qwen3-VL-8B performance by 10.5%, rivaling the capabilities of GPT-5.4 and Gemini 3.
More from Research
- Independently trained models develop 'universal' attention heads? New research shows predictable neuron scaling — CSProfKGD · 2026-08-05
- Connito Introduces Decentralized MoE Training Network on Bittensor — shibshib89 · 2026-08-05
- Improving Animal Welfare with Tech: Hyperspectral Imaging and E-Beam Vaccines in Poultry — NikoMcCarty · 2026-08-05
- Ex-Citadel Quant Releases Deep Guide on AI Power Pricing and Data Centers — PandaAshwinee · 2026-08-05
- Microsoft-Backed Pathology Foundation Model PRISM2 Published in Nature — anshulkundaje · 2026-08-05
- NeurIPS Review Reflections: Compute Barriers and Score Calibration Issues — chhaviyadav_ · 2026-08-05