LEGO: lifting-free exocentric-to-egocentric video generation beats depth-lifting SOTA pipelines
25frms · hf · 2026-10-09
- Task: generating egocentric video from a single exocentric recording—little view overlap, mostly unobserved target view.
- Unlike SOTA explicit pipelines (depth → point cloud lifting → reprojection, where depth errors become misplaced content), LEGO fine-tunes an LVSM-style view synthesizer to render the egocentric view directly, resolving cross-view correspondence internally.
- Key argument: its probabilistic mapping trades fine texture for structural alignment—ideal for a diffusion generator whose denoising restores detail; per-region confidence masks low-confidence areas and guides early denoising steps.
- Consistently outperforms the explicit SOTA and generalizes to other datasets without retraining.
More from Multimodal
- You can now spot Opus AI video slop by its sound: synced beats as a fingerprint — hudzah · 2026-10-09
- Monkey King riding a tiger: AI video nails a stunning Chinese-style action scene — lucky-plume · 2026-10-09
- Autoregressive Retriever (ARR) Refines Queries with Retrieved Item Feedback via SFT and RL — _reachsumit · 2026-10-09
- Sony's Syn-Omni: Shared + Expert LoRA Paths Beat Omnimodal Embedding Baselines Across 81 Tasks — _reachsumit · 2026-10-09
- RISEBench++: 65 reasoning-based visual editing tasks; best model GPT-Image-2.5 hits only 56.6% — VisionXLab · 2026-10-09
- VibeEdit Replaces Text Prompts with Canvas Marks, Scoring 79.9 on Edit Benchmark — Sydney-Uni · 2026-10-09