REPA may stop helping as image encoders scale up, a new technical thread argues

SeunghyunSEO7 · x · 2026-07-29

A technical discussion argues that, once models scale far enough, initializing a vision encoder from a smaller external ImageNet-trained encoder may stop being the right choice.

The poster says that although REPA has had a strong impact on image generation, the broader scaling question is whether a model should simply learn the representation from scratch instead of relying on a 0.6B–1B pretrained external encoder. The attached figure highlights two findings: as DINO models scale up, generation quality with REPA can paradoxically degrade, and the denoising signal can shift from a helpful booster to a brake over training.

Related event: Training Vision Encoders from Scratch Proves Superior with Adequate Budget(3 posts)→

Original post →

More from Research

Research channel →