EchoChange Fixes AR Errors, Beats 11 Baselines in Remote Sensing Disaster Captioning

青稞AI · wechat · 2026-08-23

Researchers from Xi'an Jiaotong University and CAS propose EchoChange, modeling bi-temporal remote sensing change captioning as a discrete mask token diffusion task. It employs a dual-pass remasking strategy to enable self-correction, addressing the error cascading issue in autoregressive models.

Key Innovations

Performance

Based on Qwen3.5-VL 9B, EchoChange achieves absolute improvements of 11.19–18.88 points in ROUGE-L, METEOR, and ST5-SCS on the RSCC benchmark, outperforming 11 baselines including Qwen2-VL and InternVL3. In factual correction tasks, it achieves a single-error correction rate of 65.34% with a clean retention rate of 79.26%, demonstrating selective revision capabilities.

Original post →

More from Multimodal

Multimodal channel →