ReactVAU at ECCV 2026: fast-slow streaming video understanding cuts heavy MLLM calls
CMHungSteven · x · 2026-09-10
Presented at ECCV 2026 in Malmö, ReactVAU is a causal framework for streaming video understanding that avoids peeking at the future and avoids waking a heavy MLLM every normal second.
Design:
- An always-on fast detector monitors every frame;
- Persistent memory captures rare cues;
- Slow LLM-based reasoning fires only when the stream looks suspicious.
Under a causal protocol it achieves competitive detection and description performance with far fewer MLLM calls. Paper, code, weights, and demo are public; a joint NVIDIA × NTHU team effort, poster Thursday Sep 10, 16:30–18:30 CEST.
More from Research
- Meta's Boxer at ECCV: closing the 3D ground-truth gap with 2D scaling — ducha_aiki · 2026-09-10
- SyncWorld turns world models into zero-shot robot simulators via visual calibration — Yuncong Yang · 2026-09-10
- Feng Yao wins ECVA PhD Award at ECCV 2026 for 3D humans + language thesis — Michael_J_Black · 2026-09-10
- Devs call for standardized "model performance across harnesses" evals — zainhas · 2026-09-10
- MIT's Point2Pose tracks unknown objects in 6D with full-occlusion recovery — joemeno · 2026-09-10
- Scientists grapple with AI agent fleets reshaping research and peer review — vykthur · 2026-09-10