Qwen-Drive-1.0-4B: Alibaba's compact 4B vision-language model for autonomous driving
solyarisoftware · x · 2026-09-06
Qwen-Drive-1.0-4B has appeared on Hugging Face: a compact image-text-to-text model built for autonomous driving. It fuses vision and language to understand road scenes, plan driving maneuvers, and answer questions about the environment rather than just processing pixels.
At only 4B parameters, the model targets resource-constrained in-vehicle scenarios and marks Qwen's extension into driving applications.
More from Embodied
- Xiaomi humanoid robot logs 98% success on factory tasks with 66 DoF, half in the hands — Olivier__OG · 2026-09-06
- Robots get their GPT-3 moment: In-Context Learning lets them learn new tasks from one demo — 量子位 · 2026-09-06
- ETH Zurich open-sources full 2026 robot learning course: 12 weeks, VLA models, free — Syntetisaattori · 2026-09-06
- Yacine Builds His Own Eye-Tracking Hardware with Custom Drivers for High-FPS Tracking — yacineMTB · 2026-09-06
- StarVLA open-sources VLAct: VLA backbone trained on 16 GPUs beats NVIDIA GR00T N1.6 with 20% of data — 机器之心 · 2026-09-06
- Show Lab sweeps 3/3 at CoRL 2026: MetaWAM hits 68.9% on RoboCasa with 39% lower inference latency — MikeShou1 · 2026-09-06