Pretraining on Europe alone beats global data across every task, 10-21 point gap in remote sensing study
anselm · x · 2026-09-12
A controlled study built on SatMAE found that pretraining a geospatial foundation model on a single continent outperforms a globally balanced dataset on every downstream task tested. The team built seven pretraining sets, each with 700,000 Sentinel-2 samples, varying only the source continent (Europe-only, Africa-only, Asia-only, etc.), plus a Global set with equal samples from six continents. Each model was finetuned on four benchmarks: FMoW scene classification, MOSAIKS population density, ForTy landcover segmentation, and the six-task GEO-Bench. Performance gaps between source continents reached 10 to 21 metric points, challenging the "more diversity is better" assumption underlying half the field's sampling strategies.
More from Research
- Clay Math Institute says the Navier-Stokes problem 'has apparently been settled' — badumtsssst · 2026-09-12
- Decagon shares 19+ ablations on using GEPA for test-driven prompt optimization in production — kastnerkyle · 2026-09-12
- GraphED: graph-based AI learns how solids deform by sharing law structure across materials — bravo_abad · 2026-09-12
- GeoGuessr as an RL env: 4B VLM trained with OpenEnv and TRL to play the game — SergioPaniego · 2026-09-12
- Blur-to-video: SIGGRAPH Asia 2025 work recovers past, present and future frames from one motion-blurred photo — CSProfKGD · 2026-09-12
- IEEE Spectrum revisits how Lotfi Zadeh defied his critics to invent fuzzy logic — ArtificialOther · 2026-09-12