ConsiSpace lifts video spatial reasoning by 12.6 points with geometry-aware memory
Ting Huang · hf · 2026-07-22
ConsiSpace is a video spatial-reasoning framework aimed at long-horizon perception tasks such as navigation and video QA, where models must keep spatial evidence consistent across changing viewpoints.
What the method does
- Introduces a geometry-consistency-aware memory that combines implicit evidence tokens with explicit geometric cues.
- Uses compact evidence organization to avoid redundant observations while preserving task-relevant spatial information.
- Adds Unified Consistency Self-Supervised Reinforcement Learning (UC-SSRL) after SFT, with rewards tied to answer, metric, and topology consistency.
Results
- Evaluated on VSI-Bench, OSI-Bench, and MMSI-Video-Bench.
- The method improves the average score by 12.6 points over the strongest baselines.
The core claim is that spatial reasoning becomes more stable when the model is trained to preserve geometric consistency explicitly, not just semantic similarity.
More from Multimodal
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11
- MiniMax Music Production Toolkit 2.5 for ComfyUI adds full mastering chain — Vivid_Promise1700 · 2026-09-11
- New Node Finder for ComfyUI ranks fresh nodes by star velocity and recency — Luke2642 · 2026-09-11
- Using a finisher move on one mosquito with MiniMax H3 MAX — the bug survives — Hailuo_AI · 2026-09-11