DroneWAM: JEPA world-action model boosts drone navigation with adaptive rollout depth
Liang Yao · hf · 2026-09-29
DroneWAM is an efficient world-action model for drone visual navigation that predicts how candidate actions change future observations.
- JEPA-based architecture models future states in representation space, avoiding costly future image generation; a pretrained Resampler compresses dense features into fewer latent tokens per imagined step
- Adaptive rollout with a preference-trained Gate allocates prediction depth per scene, cutting average depth from 8 to 4.58 while improving trajectory accuracy
- Released DroneNav-6D simulated dataset with synchronized RGB, 6-DoF trajectories, control commands, and randomized wind disturbances
Achieves best trajectory accuracy among compared methods; code and data to be open-sourced.
More from Embodied
- Shanghai AI Lab's InfiniHand Tracks 3D Hands in World Space at 11 FPS from Egocentric Video — Shanghai-AI-Laboratory · 2026-09-29
- Inside Dexory's factory: how warehouse-scanning robots go from R&D to assembly in three stages — lukas_m_ziegler · 2026-09-29
- EgoDemo, an egocentric human demonstration dataset for embodied AI, trends on HF — LightwheelAI · 2026-09-29
- Leaked OpenAI 'dot' details show raising phone to ear triggers ChatGPT Voice — koltregaskes · 2026-09-29
- Uniformation S10 Robotic Resin Printer Handles Printing, Washing and Drying Automatically — philfung · 2026-09-29
- REALM generates reactive listener facial motion, deployed on an Ameca humanoid robot — MacquarieUni · 2026-09-29