LEAP uses learned block-wise evidence retrieval to boost hour-scale audio-video QA by up to 16.8%
Juyi Lin · hf · 2026-10-01
LEAP tackles the context dilemma in hour-scale audio-visual QA: instead of encoding whole recordings or uniform compression, it splits recordings into fixed-duration blocks, scores candidate windows with a lightweight localization pass, and re-encodes only top-ranked windows in a bounded answer pass — keeping input length independent of recording duration. Localization runs on pre-computed transcripts without decoding frames, while the final pass reads raw audio-visual streams to preserve fine-grained evidence. Both stages are trained with LoRAs, and the block grid enables streaming inference without streaming-specific training. LEAP improves over the Qwen3-Omni-30B-A3B baseline by 4.5-16.8% across AVQA benchmarks and transfers to MiniCPM-o 4.5, beating its published results by 3.1-13.0%.
More from Multimodal
- A Qwen outpaint LoRA beats native outpainting on quality and consistency — linoy_tsaban · 2026-10-01
- Midjourney Artist Opens Her Entire Archive and Site mjpro.ai for Free — ciguleva · 2026-10-01
- Seeddance 2.5 used to create a 3D animated Louvre heist scene, prompt shared — azed_ai · 2026-10-01
- Orbit LoRA plus first/last frames freezes time in MiniMax video with 360 camera moves — AndrewJumpen · 2026-10-01
- Google AI turns a 500-year-old Asian lament into cinematic short film Sabalga — hansori_kr · 2026-10-01
- Nvidia open-sources Lyra 2.0, turning any image into an explorable 3D world — CurieuxExplorer · 2026-10-01