LEAP uses learned block-wise evidence retrieval to boost hour-scale audio-video QA by up to 16.8%

Juyi Lin · hf · 2026-10-01

LEAP tackles the context dilemma in hour-scale audio-visual QA: instead of encoding whole recordings or uniform compression, it splits recordings into fixed-duration blocks, scores candidate windows with a lightweight localization pass, and re-encodes only top-ranked windows in a bounded answer pass — keeping input length independent of recording duration. Localization runs on pre-computed transcripts without decoding frames, while the final pass reads raw audio-visual streams to preserve fine-grained evidence. Both stages are trained with LoRAs, and the block grid enables streaming inference without streaming-specific training. LEAP improves over the Qwen3-Omni-30B-A3B baseline by 4.5-16.8% across AVQA benchmarks and transfers to MiniCPM-o 4.5, beating its published results by 3.1-13.0%.

Original post →

More from Multimodal

Multimodal channel →