FairRSFM benchmark exposes biome-level bias in remote sensing foundation models
Md Aminur Hossain · hf · 2026-10-07
FairRSFM is a biome-aware benchmark and debiasing framework showing that aggregate metrics mask systematic performance disparities in remote sensing foundation models.
- It maps 14 terrestrial biomes into six ecological macro-groups and evaluates models under a frozen-backbone protocol across four downstream datasets.
- Prithvi-EO-2.0 scores 90.98% overall macro-F1 on m-EuroSAT but only 83.72% mean worst-group; m-SA-Crop-Type mIoU drops from 27.30% to 18.47% in the Xeric/Mineralogical group.
- Three mitigation baselines (BOLP, DBR, GroupDRO) are tested; BOLP lifts Prithvi-EO-2.0 worst-group F1@opt on m-BigEarthNet from 46.12% to 50.27% without touching the backbone.
- Code and datasets are open-sourced.
More from Research
- CUAWright: Terminal-Only Computer-Use Agent Beats GUI Harnesses, Cuts Cost 37.5% — ysu_nlp · 2026-10-07
- AI's Top 10 papers list: Rulin Shao lands two first-author picks — ShayneRedford · 2026-10-07
- Paradigm evals its math model across 7 hard benchmarks, releases full eval suite — tensorqt · 2026-10-07
- Paradigm: post-training gains hinge on combining procedural and LLM-based synthetic data — tensorqt · 2026-10-07
- Paradigm scales RL context from 65k to 131k tokens using a trained value model — tensorqt · 2026-10-07
- Limite 1B borrows nanogpt speedrun architecture: NorMuon, MUDD variant and XSA — tensorqt · 2026-10-07