EM²Mem: event-centric multimodal memory for long-video QA
zjunlp · hf · 2026-09-02
ZJUNLP proposes EM²Mem, which binds multimodal evidence to event anchors for compact, generation-ready memory, improving long-video question answering.
More from Research
- OpenAI's 'Recurrent Depth' Reasoning Raises Monitoring Concerns — GaryMarcus · 2026-09-02
- Internal Activation Loops vs. Token Conversion in CoT Reasoning — gandamu_ml · 2026-09-02
- Hidden trade-off in video world models: geometry vs scale — keenanisalive · 2026-09-02
- Analyst report massively overestimates robot data generation — zephyr_z9 · 2026-09-02
- You can distill consistent surface meshes from the Atlas world model — MatthewChang · 2026-09-02
- NoRA: Normalized LoRA Boosts Convergence and Stability — burny_tech · 2026-09-02