GEAR Uses Geometry as Address to Route Attention, Enabling Minute-Long Camera-Controlled Video Generation

Zesong Yang · hf · 2026-09-30

The GEAR paper on Hugging Face introduces a Geometry-Enabled Attention Routing framework for long-horizon camera-controlled video generation. Key insight: geometry need not explain the scene — it only determines where to read from visual memory, while attention decides what to recover. GEAR keeps frame latents (avoiding error-prone global 3D fusion), uses Geometric Correspondence Attention to inject geometrically matched historical features during denoising, and adds an Invisible Octree to reject occluded correspondences. It achieves SOTA visual quality, precise camera control, and revisit consistency on minute-long trajectories.

Original post →

More from Multimodal

Multimodal channel →