EchoChange Fixes AR Errors, Beats 11 Baselines in Remote Sensing Disaster Captioning
青稞AI · wechat · 2026-08-23
Researchers from Xi'an Jiaotong University and CAS propose EchoChange, modeling bi-temporal remote sensing change captioning as a discrete mask token diffusion task. It employs a dual-pass remasking strategy to enable self-correction, addressing the error cascading issue in autoregressive models.
Key Innovations
- Discrete Diffusion over Autoregression: Iteratively denoise from fully masked answers, allowing previously generated tokens to be remasked and revised.
- Dual-Pass Remasking: The first pass drafts from corrupted ground truth; the second pass remasks low-confidence positions, forcing the model to recover from its own erroneous draft.
- Curriculum Sampling & Calibration: Training transitions from filling few gaps to full masking; inference uses cross-step stability for remasking decisions.
Performance
Based on Qwen3.5-VL 9B, EchoChange achieves absolute improvements of 11.19–18.88 points in ROUGE-L, METEOR, and ST5-SCS on the RSCC benchmark, outperforming 11 baselines including Qwen2-VL and InternVL3. In factual correction tasks, it achieves a single-error correction rate of 65.34% with a clean retention rate of 79.26%, demonstrating selective revision capabilities.
More from Multimodal
- NAPE Audio Pretraining Achieves SOTA Without Decoders — kastnerkyle · 2026-08-24
- Local AI Generation: Pudgy Penguins Music Video with MiniMax H3 — cocktailpeanut · 2026-08-24
- Thrixel generates playable 3D forest and boss in one hour — RanaHanocka · 2026-08-24
- 30-year animator brings characters to life with AI in "The Latent Stroke" — Paferkas-71 · 2026-08-24
- D&D world opening cinematic made with Minimax H3, full workflow shared — Routine_Ad_3391 · 2026-08-24
- Trained two 16M-param models to do generative CAD with real physics — debreuil · 2026-08-24