PixelDense: Dense-Prediction Teachers Beat Semantic Encoders for Diffusion REPA, GenEval Hits 0.8093
Lehan Yang · hf · 2026-10-02
PixelDense challenges the convention that REPA alignment targets should be semantic encoders like DINOv2 or CLIP: dense-prediction foundation models (SAM2, Depth Anything v2, Metric3D v2) each outperform the DINOv2-only baseline in pixel-space diffusion, but a flat sum of all four teachers underperforms the best single geometric teacher as semantic and geometric gradients compete for one denoiser projection. PixelDense routes DINOv2+SAM2 through a semantic projection stream and Depth Anything v2+Metric3D v2 through a geometric stream with a weight-space orthogonality penalty, freezing teachers during training and dropping them at inference. Applied to PixelGen and DeCo with one recipe, it raises PixelGen-XXL's GenEval Overall from 0.7927 to 0.8093, achieves up to 53.1% PQ gain and 36.0% depth AbsRel reduction, converges 1.23x faster, and improves SDEdit background PSNR by up to 2.2 dB.
More from Multimodal
- Image editing demo with Nano Banana — tkasasagi · 2026-10-02
- Creator: Opus 5.5 edits videos well, but forcing AI to clip without real need yields garbage — AlchainHust · 2026-10-02
- Reddit user's VEC concept car AI video shows startlingly realistic motion — Vashukanni · 2026-10-02
- Fable 5.5 rumored: webpage morphs its art style to match passing images — lxfater · 2026-10-02
- Oxford VGG Unveils SynCity 3000, Generating Globally Coherent Scene-Scale 3D Worlds — rsasaki0109 · 2026-10-02
- Omni-Embed-Mini: A 0.9B Embedder Adds Five Modalities Without Touching Text Weights — _reachsumit · 2026-10-02