FixAnything: 20 paired videos of LoRA let a video diffusion model repair any 3D rendering
kwangmoo_yi · x · 2026-08-26
CMU's FixAnything (ECCV 2026) uses one generalist video model to repair rendering artifacts from any 3D representation:
- Approach: framed as video-to-video translation. 3DGS/NeRF/mesh/point-cloud renders each have artifacts, but a render still preserves camera trajectory and coarse layout — an effective control signal for a pretrained video diffusion model (Wan2.1-I2V-14B)
- Data efficiency: even a very sparse point cloud (e.g. from COLMAP) gives effective camera control, with minimal LoRA finetuning on 20 paired videos
- Alignment: camera poses recovered by COLMAP from outputs serve as the reward; Flow-DPO steers toward geometrically consistent results
- Anchoring: clean training views along the trajectory act as anchors, propagating appearance, lighting and structure into degraded in-between frames
Related event: FixAnything Unifies 3D Rendering Artifact Repair with Video Priors(2 posts)→
More from Multimodal
- LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training — Andreas Hochlehnert · 2026-08-26
- Redditor shares AI-generated anime series 'Realmz of the Redeemers' ep 2 — Kindly_Poet_7878 · 2026-08-26
- Captain America x Harry Potter AI video on a single RTX 3070: 10s clip in ~10 minutes — luka06111 · 2026-08-26
- Redditor shares dark fantasy AI-generated video clip — Illustrious_Fee_3676 · 2026-08-26
- Claude-built prompt templating system mass-generates ComfyUI workflows; 5-min video re-renders in 2 hours — spikyness27 · 2026-08-26
- Qwen Image Edit appears trained at 1MP: resizing inputs to 1024x1024 yields far better results — Civil_Fee_7862 · 2026-08-26