PixelDense: Dense-Prediction Teachers Beat Semantic Encoders for Diffusion REPA, GenEval Hits 0.8093

Lehan Yang · hf · 2026-10-02

PixelDense challenges the convention that REPA alignment targets should be semantic encoders like DINOv2 or CLIP: dense-prediction foundation models (SAM2, Depth Anything v2, Metric3D v2) each outperform the DINOv2-only baseline in pixel-space diffusion, but a flat sum of all four teachers underperforms the best single geometric teacher as semantic and geometric gradients compete for one denoiser projection. PixelDense routes DINOv2+SAM2 through a semantic projection stream and Depth Anything v2+Metric3D v2 through a geometric stream with a weight-space orthogonality penalty, freezing teachers during training and dropping them at inference. Applied to PixelGen and DeCo with one recipe, it raises PixelGen-XXL's GenEval Overall from 0.7927 to 0.8093, achieves up to 53.1% PQ gain and 36.0% depth AbsRel reduction, converges 1.23x faster, and improves SDEdit background PSNR by up to 2.2 dB.

Original post →

More from Multimodal

Multimodal channel →