Why block masks work in JEPA: 151 pretraining runs show unrecoverable content is key

udmrzn · x · 2026-10-01

An arXiv paper gives a measurement account of JEPA's masking sensitivity: block masks work not because of their shape but because they hide coarse-scale content that low-level interpolation cannot recover. Across 151 pretraining runs on ImageNet-100, strip masks matching blocks in area and contiguity hit 40.3% linear top-1 vs 64.3% for blocks; within one geometry family placements leaving the least unrecoverable content lose 6.5 points; pixel targets span 7 points where latent targets span 25; and against a frozen target the 19-point gap between random and block masks closes to 1.5. On UCF101, masking ratio decides whether content unrecoverability or context reachability binds.

Original post →

More from Research

Research channel →