DreamTraj: Predicting 6-DoF Trajectories by Reading Unrendered Video Diffusion Latents
SUSTech · hf · 2026-08-04
The paper proposes the DreamTraj model and the accompanying MOVE dataset, leveraging the internal representations of video diffusion models for efficient robotic trajectory prediction.
- Innovation: Unlike costly pipelines that rely on generating full videos before extracting motion, DreamTraj requires only a single RGB image and a task instruction. It reads motion directly from the internal representations of a frozen image-to-video diffusion model at an early denoising step.
- MOVE Dataset: Contains 5,038 egocentric object trajectories paired with fine-grained natural language instructions, filling the gap for fine-grained language-to-motion annotations.
- Performance: A lightweight flow-matching Reader decodes attention tracks into relative 6-DoF poses. Against forecasters consuming multi-frame or privileged inputs, DreamTraj sets a new SOTA on both translation and rotation, running 4.6x faster than generate-then-extract pipelines.
More from Embodied
- Testing GPT-5.6 and Gemini Robotics Models: Impressive but Need Real Deployment Data — m_wulfmeier · 2026-08-04
- Musk Says Neuralink Vision Implants Could Cure Blindness in 6-12 Months — saibharadwaj · 2026-08-04
- Tsinghua Team's Embodied AI Startup Poke Raises $100M+ in Pre-A Round — 创业邦 · 2026-08-04
- Minimax H3 Local Test: 2MP Video Generation in 51 Minutes on RTX 6000 Pro — JahJedi · 2026-08-04
- LeRobot now supports 30+ robot hardware integrations with drop-in plugins — m_wulfmeier · 2026-08-04
- NUS Creates Octopus-Inspired Swimming Robot Driven by Just Two Motors — lukas_m_ziegler · 2026-08-04