ZJU Releases EMem-Bench: 2,554 Episodes to Test Embodied Agent Memory
OmniAI-ZJU · hf · 2026-09-29
ZJU's OmniAI team introduces EmbodiedMemory-Bench (EMem-Bench), a benchmark directly measuring memory during long-horizon embodied interaction, plus a matching memory system.
- Motivation: agents must retain and update environment information across observation, action, and change, but current models show four deficiencies—weak fine-grained visual memory, unreliable dynamic world-state tracking, failure to record interaction-revealed state, and limited generalization from experience. Existing benchmarks don't test memory directly in long-horizon interaction.
- Benchmark: 2,554 interactive episodes across four task families; agents must build and update memory from interaction history, then use it to complete later tasks.
- Method: Embodied-Memorizer (EMem) organizes embodied experience into spatial, event, and scene memories; a trained 8B policy (EMem-8B) manages and uses them.
- Results: evaluations of open-source and proprietary MLLMs and multimodal memory systems show current models remain weak and uneven on the four challenges; EMem is the best overall memory system on matched backbones, and EMem-8B further improves over its backbone.
More from Embodied
- Mark Cuban predicts humanoid robots will fail in 5–10 years, homes will adapt to task-specific robots — Polymarket · 2026-09-29
- Humanoid robot reportedly in low-volume production: 5 units planned for 2026 — teortaxesTex · 2026-09-29
- Hooking a model up to a robot arm: two USB cameras, an SO-101, ~2 decisions per second — ai · 2026-09-29
- Scoble: lab AI glasses get fold-out camera covers that light up when shooting — Scobleizer · 2026-09-29
- AIST AI Center in Tokyo recruits RAs and postdocs in multi-robot coordination, up to ¥2,100/hour — HirokatuKataoka · 2026-09-29
- Edge AI chip startup SiMa.ai raises $150M Series C at $1.45B valuation — dauber · 2026-09-29