ConsiSpace lifts video spatial reasoning by 12.6 points with geometry-aware memory
Ting Huang · hf · 2026-07-22
ConsiSpace is a video spatial-reasoning framework aimed at long-horizon perception tasks such as navigation and video QA, where models must keep spatial evidence consistent across changing viewpoints.
What the method does
- Introduces a geometry-consistency-aware memory that combines implicit evidence tokens with explicit geometric cues.
- Uses compact evidence organization to avoid redundant observations while preserving task-relevant spatial information.
- Adds Unified Consistency Self-Supervised Reinforcement Learning (UC-SSRL) after SFT, with rewards tied to answer, metric, and topology consistency.
Results
- Evaluated on VSI-Bench, OSI-Bench, and MMSI-Video-Bench.
- The method improves the average score by 12.6 points over the strongest baselines.
The core claim is that spatial reasoning becomes more stable when the model is trained to preserve geometric consistency explicitly, not just semantic similarity.
More from Multimodal
- AI video’s bottleneck is workflow, not generation quality — aftahi_ai · 2026-07-22
- VideoChat3 debuts as a fully open 4B video model for long and streaming clips — 机器之心 · 2026-07-22
- GPT Image 2 Prompt Shows How to Build Editorial Posters Inside ChatGPT — SimplyAnnisa · 2026-07-22
- AlayaWorld open-sources a 720p, 24 FPS video world model with camera control — fruesome · 2026-07-22
- APOB turns AI influencers into a surprisingly smooth TikTok dance clip — aftahi_ai · 2026-07-22
- Skywork Video pitches an all-in-one AI video workspace with storyboard control — alifcoder · 2026-07-22