Qwen-Drive-1.0: a VLM foundation model unifying 3D perception and motion planning
tarantulae · x · 2026-09-02
Qwen's new arXiv paper presents Qwen-Drive-1.0, a step toward a vision-language foundation model for autonomous driving. It retains the pretrained VLM architecture, adds an external BEV head for 3D detection, occupancy and map segmentation, and a Planning Expert for trajectory generation. Staged training mixes driving supervision with general VL data, showing competitive open/closed-loop planning while preserving general VL capability.
Related event: Alibaba Releases Qwen-Drive-1.0: Unified Perception and Planning VLM(4 posts)→
More from Embodied
- Agent designs its own hardware: Astra produces USB-powered 4-light PCB files — paraschopra · 2026-09-21
- Starlink enables West Africa's first telesurgical kidney removal across 500 km — elonmusk · 2026-09-21
- RPent open-sourced: 92.6% on LIBERO-PRO and 7x faster embodied agent execution — 量子位 · 2026-09-21
- ZuckOff App Detects Nearby Meta Smart Glasses via Bluetooth Fingerprints, Tops 5,000 Downloads — LexiLove · 2026-09-21
- GaME (CVPR 2026): Gaussian mapping that forgets stale geometry as robot scenes change — lucacarlone1 · 2026-09-21
- Stanford RL method fixes VLA latency, lifting robot success from 42% to 97% with 10 minutes of data — burny_tech · 2026-09-21