GEAR Uses Geometry as Address to Route Attention, Enabling Minute-Long Camera-Controlled Video Generation
Zesong Yang · hf · 2026-09-30
The GEAR paper on Hugging Face introduces a Geometry-Enabled Attention Routing framework for long-horizon camera-controlled video generation. Key insight: geometry need not explain the scene — it only determines where to read from visual memory, while attention decides what to recover. GEAR keeps frame latents (avoiding error-prone global 3D fusion), uses Geometric Correspondence Attention to inject geometrically matched historical features during denoising, and adds an Invisible Octree to reject occluded correspondences. It achieves SOTA visual quality, precise camera control, and revisit consistency on minute-long trajectories.
More from Multimodal
- Opus 5.5 + ElevenLabs Turn a Treasury Figure Into Rock and Rap Tracks — leveragedupside · 2026-09-30
- Lumibelle: open-source video editor feeds reusable character reels into MiniMax H3 for consistency — rad_reverbererations · 2026-09-30
- Asking for help: outpainting 9:16 video to 16:9 with Minimax H3 changes the footage entirely — navarisun · 2026-09-30
- Inception Labs launches Mercury Voice, a diffusion LLM with 2x lower latency than GPT-6 Luna and Claude Haiku 4.5 — StefanoErmon · 2026-09-30
- ComfyUI-Continuity Update Adds Cast Management, RefMod Presets and Seam Optimization — Okims_kor · 2026-09-30
- AI turns its subagents into a rock band, full MV made with local music and video models — WolframRvnwlf · 2026-09-30