Alibaba's Qwen-Drive: Autonomous Driving Model Preserving VLM Architecture
aigclink · x · 2026-09-02
Alibaba released Qwen-Drive 1.0, an autonomous driving foundation model that retains the full architecture of a pre-trained VLM instead of using an end-to-end specialized model. It achieves driving capabilities via external modules.
Technical Architecture:
- BEV Perception Head: Converts visuals to 3D road conditions, handling 3D object detection, semantic occupancy prediction, and BEV map segmentation.
- Planning Expert: Generates a 5-second trajectory (50 waypoints) based on visual input using flow matching for smooth, safe routing.
Performance: Achieves a score of 90.7 on the NAVSIM trajectory planning benchmark, comparable to SOTA professional planners.
Related event: Alibaba Releases Qwen-Drive-1.0: Unified Perception and Planning VLM(4 posts)→
More from Embodied
- Overlaying object tracking on OpenStreetMap with GNSS/IMU trajectories and depth estimation — rsasaki0109 · 2026-09-21
- Asking for the best way to run MiniMax H3 video generation on a DGX Spark — Shady-Dragon · 2026-09-21
- Meta reportedly building camera-free smart glasses codenamed Luna for privacy — emmanuelvivier · 2026-09-21
- AgiBot tops global humanoid shipments in H1; Chinese firms hold over 97% of market — emmanuelvivier · 2026-09-21
- Chinese T800 humanoid patrols Shenzhen streets with SWAT, priced from $25,000 — TinfoilTricorn · 2026-09-21
- Agent designs its own hardware: Astra produces USB-powered 4-light PCB files — paraschopra · 2026-09-21