Blurry renders beat sharp point clouds: LEGO hits 20.28 dB PSNR on Ego-Exo4D

LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation

Suhwan Cho, Yonwoo Choi, Soongjin Kim, Jicheol Park, Taegyu Lim

cs.CV

2026-10-09

LEGO fine-tunes an LVSM to render the headset view without depth or point clouds, then conditions Wan on that soft image. Seen-split PSNR is 20.28 versus EgoX rerun at 15.57.

What problem this solves

First-person video is scarce. Third-person footage is not. Given one exocentric clip and a headset trajectory, the task is to synthesize what that headset would record. The cameras barely overlap, so hands and nearby objects are often missing from the source.

EgoX, the strongest explicit baseline, estimates depth, lifts frames into a point cloud, re-renders along the egocentric trajectory, and feeds that image to Wan. A depth-derived attention bias is applied in training and at inference. When depth is wrong, content lands in the wrong place and stays sharp, with no mark that it is untrustworthy. Denoising is good at sharpening blur. Putting an object back where it belongs is a harder fix.

Method

LEGO does not change the generator. It changes the condition. The view synthesizer is an LVSM-style transformer: 24 blocks, width 768, 12 heads, patch size 8. Geometry enters only as Plücker rays, a direction plus the cross product of the camera center and that direction. Poses are written in the exocentric frame, so the cross product carries the baseline between the cameras. One exocentric frame and two ray maps go in. The egocentric frame comes out. The loss is ℓ2 plus 0.5 times LPIPS, with no depth labels and no correspondence labels.

The public LVSM checkpoint only interpolates nearby pinhole views. Here the target is an Aria fisheye with little overlap. Fisheye rays are recovered per pixel, then all weights are fine-tuned on Ego-Exo4D for 10,000 steps, about 3.5 hours on four H200s, and frozen. Before fine-tuning, 95.9% of the area outside the lens circle is filled with scene content. After, spill is 0.1%.

Attention from the last 8 layers, averaged over heads, is a distribution over source locations for each target pixel. When that distribution spreads, the pixel is an average of several candidates: texture washes out, placement stays nearby. Confidence is the probability mass of the top 16 sources. Locations under 0.3 are painted neutral gray, which the VAE encodes near an empty condition. A point cloud can leave a hole where geometry is missing. This average fills unseen regions anyway, so the gate is how it abstains.

The backbone is Wan2.1-I2V-14B with LoRA rank 256, on a 49×448×1232 canvas with the two views side by side. The condition sits in the egocentric half. Training runs 20,000 steps. Inference uses 50 steps, guidance 5, and a text prompt. Renders and confidence maps are precomputed per clip.

Sampling is flow matching. Layout is set in the early high-noise steps. For the first 80% of steps, ARC pulls the current clean estimate toward the render by the square of confidence. High-confidence regions move a lot. Low-confidence regions barely move. Later steps add texture the render does not contain. No extra training. The 80% cutoff and the square were chosen on the unseen split, because there is no validation split.

Results

Ego-Exo4D has seen and unseen scenes. Published EgoX numbers sit beside EgoX†, the released checkpoint scored with this paper's code. The rerun is lower: seen PSNR falls from 16.05 to 15.57, unseen from 14.38 to 13.53.

SplitMethodPSNR↑LPIPS↓FVD↓
seenEgoX / EgoX†16.05 / 15.570.498 / 0.466184.47 / 197.95
seenLEGO20.280.310155.56
unseenEgoX / EgoX†14.38 / 13.530.552 / 0.542440.64 / 455.06
unseenLEGO15.910.504404.77

Seen SSIM is 0.689 against 0.556, and CLIP-I is 0.925 against the published 0.896. Unseen SSIM is 0.505 against 0.457, CLIP-I 0.897 against 0.877. Flicker, smoothness, and dynamic degree are already saturated. Exo2Ego-V reaches only 14.53 PSNR on seen scenes. TrajectoryCrafter, Wan Fun Control, and Wan VACE are lower.

Under one recipe, seen PSNR is 12.30 with an empty condition, 15.59 with the point cloud, 15.71 with the untuned synthesizer, and 18.24 after fine-tuning. The untuned model matches the point cloud, so the gain is the fisheye adaptation. Gating leaves PSNR at 18.27 and cuts FVD from 183.40 to 166.86. Restoring exocentric tokens reaches 18.91 and 159.04. Adding geometry-guided attention (GGA) on top drops PSNR to 18.62 and raises FVD to 194.98. ARC anchored on the point cloud falls to 13.96. Anchored on the synthesizer, it reaches 20.28 and FVD 155.56.

On unseen scenes the synthesizer render alone is not uniformly better: PSNR 14.35 against 13.42, but FVD 581.90 against 549.77. Gating, the source view, and ARC together bring FVD to 404.77.

The blurrier render is the better aligned one. On unseen scenes, ZNCC at patch sizes 16/32/64 is 0.169/0.217/0.275 for the synthesizer and 0.034/0.047/0.059 for the point cloud. DINO similarity is 0.586 against 0.476. Sharpness relative to the recording is 0.28 against 0.97. Alignment inside the valid circle is about 4.2 dB higher, and coverage is 98.1% against 53.4%. Output sharpness comes back, with ratios 1.04 and 1.10. Color does not. Chroma ratio is 0.41 in the render and 0.45 in the output, against 0.82 for a point-cloud condition and 0.62 after gating. Where correspondence is uncertain, color is averaged toward gray.

Confidence tracks error. At a threshold of 0.3, luminance error of kept versus discarded windows is 11.6 against 16.6 on seen scenes and 25.9 against 30.7 on unseen scenes. Output error follows: 13.1 against 21.4, and 26.1 against 32.8.

With no further training, 200 clips each from EgoHumans and Nymeria score PSNR 15.52 against 13.74 and 15.23 against 11.64, and FVD 332.41 against 421.83 and 210.27 against 296.98, all versus EgoX†. The gate keeps about 15% of the render, close to unseen Ego-Exo4D.

Object boxes barely move. Under a protocol that matches the published unseen numbers, LocErr is 143.3 against 146.5 and box IoU is 0.097 against 0.091. On seen scenes the same protocol scores LEGO at 107.7 and the released checkpoint at 125.3, both far from the published 61.8.

Building the condition drops from 69.2 seconds to 4.7. End to end on one H200 is 577.9 seconds against 666.7, at 68.2 GiB against 69.9. The denoiser still takes most of the time.

Why it matters

Both poses have to be known, and the headset has to be an Aria-style fisheye. That is where depth misplaces hands and nearby objects. The swap is the condition. The Wan and LoRA setup can stay, and the synthesizer can be replaced later.

About 4.7 dB on seen scenes versus the rerun is a clear gap. On unseen scenes it is about 1.5 dB versus the published table and about 2.4 dB versus the rerun, the size of a condition redesign. Temporal scores were already saturated. What moves is how tightly frames match the recording. A sharp but misplaced point cloud gets worse when ARC pulls toward it.

Limitations

The synthesizer runs per frame on a single exocentric image. Regions the source never saw are an empty condition, filled from the diffusion prior. Larger viewpoint gaps make that hole harder to ignore.

Color drifts toward gray on unfamiliar scenes. Unseen outputs have a chroma ratio of 0.45, against 0.82 with a point-cloud condition. FVD is sensitive to per-frame appearance, so part of the unseen FVD gain is color.

The 80% schedule and the squared confidence were selected on the unseen split. With no validation split, that row includes a hyperparameter choice.

Compute does not match. LEGO trains about four days on four H200s. EgoX trains about one day on eight H200s, with the same per-GPU batch of 1. Learning rate and step count line up. Total GPU time does not. Metrics were reimplemented, so comparisons with the original table should use EgoX†.

The gate keeps roughly 15% of the render. The rest is gray, and the generator can still see the exocentric tokens. How much structure comes from the synthesizer is not isolated. Box error barely changes, so the impression that hands land correctly is not backed by the object metrics.

Poses are required, and noisy poses are not tested. In-the-wild clips have no ground truth. Prompts for EgoHumans and Nymeria were written by Qwen2.5-VL and shared by both methods, so the ranking is fair and the absolute scores are not comparable to EgoX's original prompts.

Terms

Source

Related papers

All paper explainers