ZJU's Spatial-Interactor teaches VLMs spatial reasoning via physical interaction
OmniAI-ZJU · hf · 2026-09-24
- Targets a core VLM weakness: perceiving local state transitions from object motion and viewpoint changes, and integrating them over long trajectories.
- Spatial-Interactor trains VLMs through interaction with the observable physical world via a three-level curriculum: L1 passive world-state transitions, L2 active self-state transitions, L3 long-horizon interaction trajectories.
- Introduces the LSI-108K dataset built from simulated and real interaction trajectories, with tasks aligned to each level.
- Two-stage training: SFT for L1/L2 local transition modeling, then On-Policy Distillation where a privileged teacher branch (given segment-level transition descriptions) supervises the student's on-policy CoT for long-horizon integration.
- Consistent gains across multiple VLMs and spatial benchmarks.
More from Research
- Dual-H200 fine-tune of Marigold V2 with 16-frame temporal attention aims to fix video depth flicker — AntonObukhov1 · 2026-09-24
- Same Prompt, Opposite Results: GPT-4 Goes Silent 30/30 Where GPT-3.5 Never Stops — rayanpal_ · 2026-09-24
- Amazon's BoundaryMORPH uses Gaussian Processes to budget cross-encoder reranking, +5.4 nCG@100 — _reachsumit · 2026-09-24
- Paper: retrieval recall ceilings LLM recommendation reranking — oracle eval inflates NDCG up to 95%, none beat CF — _reachsumit · 2026-09-24
- Beyond a scalar: distributional serving interfaces let multiple task heads reuse watch-time distributions — _reachsumit · 2026-09-24
- Meta's Muse Realtime Avatar beats two leading commercial avatar systems in blind tests — AIatMeta · 2026-09-24