Qwen-Drive-1.0: a VLM foundation model unifying 3D perception and motion planning

tarantulae · x · 2026-09-02

Qwen's new arXiv paper presents Qwen-Drive-1.0, a step toward a vision-language foundation model for autonomous driving. It retains the pretrained VLM architecture, adds an external BEV head for 3D detection, occupancy and map segmentation, and a Planning Expert for trajectory generation. Staged training mixes driving supervision with general VL data, showing competitive open/closed-loop planning while preserving general VL capability.

Related event: Alibaba Releases Qwen-Drive-1.0: Unified Perception and Planning VLM(4 posts)→

Original post →

More from Embodied

Embodied channel →