REPA may stop helping as image encoders scale up, a new technical thread argues
SeunghyunSEO7 · x · 2026-07-29
A technical discussion argues that, once models scale far enough, initializing a vision encoder from a smaller external ImageNet-trained encoder may stop being the right choice.
The poster says that although REPA has had a strong impact on image generation, the broader scaling question is whether a model should simply learn the representation from scratch instead of relying on a 0.6B–1B pretrained external encoder. The attached figure highlights two findings: as DINO models scale up, generation quality with REPA can paradoxically degrade, and the denoising signal can shift from a helpful booster to a brake over training.
Related event: Training Vision Encoders from Scratch Proves Superior with Adequate Budget(3 posts)→
More from Research
- Discussion: Why hasn't anyone built a neural network to detect AI text? Image detection has research papers — emeka_boris · 2026-07-30
- Agents still struggle with mathematical work: Codex spirals into 'proof certificates' and inventories — doodlestein · 2026-07-30
- TorchSpec Enables Disaggregated Speculative Decoding Training at Scale — zhyncs42 · 2026-07-30
- Compute Surge: 10 Major Scientific Breakthroughs AI Could Unlock by 2028 — Annual_Judge_7272 · 2026-07-30
- Inside SOTA Deep Research: Native Model Training and 150 Sub-Agents — SimonShaoleiDu · 2026-07-30
- Princeton Prof Reviews MIT Nonconvex Optimization Paper: Prize Remains Open — HazanPrinceton · 2026-07-30