ReflectWorld-MM stores video memory around entities, not frames, and tops 6 benchmarks
Xiaokang Ma · hf · 2026-07-21
ReflectWorld-MM proposes an entity-oriented multimodal memory system for open-ended video streams. - It addresses the limits of existing video-memory systems, which usually store memories in the model context or a flat feature store and organize them around frames rather than persistent entities. - The system combines three components: a perception front end that converts audiovisual streams into entity-resolved observations, a hierarchical long-term memory inspired by human memory theory, and a full implementation that can ingest arbitrary streams and plug into off-the-shelf assistants. - Its long-term memory includes multi-scale episodic memory, evolving entity-centric semantic memory, and procedural memory. - Across six long-video and lifelong-memory benchmarks, the paper reports best accuracy on all six, beating strong memory agents and a frontier model.
More from Multimodal
- Seedance 2.0 turns one reference image into a cinematic fight scene — techhalla · 2026-07-21
- Seedance 2.0 keeps character consistency across 15+ shots with just 3 prompts — techhalla · 2026-07-21
- Google Gemini’s Omni text-to-video output is getting better, user says — michaelrabone · 2026-07-21
- OpenArt AI demos a Video Remix tool that can transform an existing video — eyishazyer · 2026-07-21
- ElevenLabs raises ElevenMusic free usage to 400 tracks a month — lukeharries · 2026-07-21
- Google Gemini now watermarks every AI video it generates — Sure_Belt9076 · 2026-07-21