NJU's OneStreamer: 4B Streaming Video LLM Tops All Eight Streaming Benchmarks
NJU · hf · 2026-10-02
NJU researchers present OneStreamer, a streaming video LLM unifying perception, memory, and proactive response. Key components:
- Proactive Hierarchical Caption Memory (PHCM): produces time-grounded detail captions and event summaries; at inference, model-generated records replace revisiting historical visual features.
- Proactive State Transition Learning (PSTL): supervises only 27.5% of annotated state tokens yet outperforms dense supervision.
- OneStreamer-1M: a synthesis pipeline yielding 1M+ streaming interaction records.
The 4B model achieves the best results across all eight streaming video understanding benchmarks, and ablations show retaining generated captions improves historical QA without degrading real-time perception.
More from Multimodal
- LTX 2.5 generates 60-second clips on a 32GB RTX 5090 as software optimization beats VRAM upgrades — OpenEffect3955 · 2026-10-02
- AI-generated J-pop music video "You & Me (君と僕)" released — SeaTransportation457 · 2026-10-02
- AI music video "Sayonara, Tomo yo" (LET GO) shared — PromptDoStudio · 2026-10-02
- USTC's PhysVista benchmark exposes wide gap between VLM visual recognition and physical understanding — ustc · 2026-10-02
- Commercial-grade AI video needs ~10 rounds of feedback, not manual edits — AlchainHust · 2026-10-02
- Suno launches Speech in public beta, generating voiceovers with matching music — The Verge AI · 2026-10-02