ConsiSpace lifts video spatial reasoning by 12.6 points with geometry-aware memory
Ting Huang · hf · 2026-07-22
ConsiSpace is a video spatial-reasoning framework aimed at long-horizon perception tasks such as navigation and video QA, where models must keep spatial evidence consistent across changing viewpoints.
What the method does
- Introduces a geometry-consistency-aware memory that combines implicit evidence tokens with explicit geometric cues.
- Uses compact evidence organization to avoid redundant observations while preserving task-relevant spatial information.
- Adds Unified Consistency Self-Supervised Reinforcement Learning (UC-SSRL) after SFT, with rewards tied to answer, metric, and topology consistency.
Results
- Evaluated on VSI-Bench, OSI-Bench, and MMSI-Video-Bench.
- The method improves the average score by 12.6 points over the strongest baselines.
The core claim is that spatial reasoning becomes more stable when the model is trained to preserve geometric consistency explicitly, not just semantic similarity.
More from Multimodal
- Lumara AI Film Festival Comes to NYC Oct 26, Top AI Filmmakers to Compete — 0xAllen_ · 2026-09-11
- Pterodactyl Detective: An AI-Generated Proof-of-Concept Trailer — PterodactylDetective · 2026-09-11
- Imperium Game Trailer Showcases AI Video Generation — keaslenyt · 2026-09-11
- FLUX.2 Klein Drifts Hard on Character Expressions While Free Gemini Holds Likeness — wacomlover · 2026-09-11
- Tencent Hunyuan releases AuK code and weights on GitHub with ComfyUI and fine-tuning support — aigclink · 2026-09-11
- Creator turns Bahamut vs Tiamat rivalry into an AI cinematic battle with Midjourney, GPT Image 2 and Seedance — azed_ai · 2026-09-11