Open-Source 10B VLX-Seek-1.5 Model for Embodied Visual Grounding Released
pmttyji · reddit · 2026-08-10
omlab released VLX-Seek-1.5-10B on Hugging Face, an open-source 10B model for fine-grained perception and visual grounding in embodied scenarios. It uses region retrieval and reference instead of coordinate generation, targeting drones, robots, surveillance, and edge devices. Highlights include stronger visual capability, faster inference, and explicit absent-target rejection. Sizes: 0.6B, 3B, 10B.
More from Embodied
- Canadian Team Wins DARPA Challenge: 29lb Drone Carries 112lb Payload — wavefnx · 2026-08-10
- NeurIPS 2026 Calls for Papers on Physical Understanding for Embodied AI — shaohua0116 · 2026-08-10
- Former Li Auto ADAS chief Lang Xianpeng raises hundreds of millions in 90 days for embodied AI startup Kunlun Xing, says training costs billions — 量子位 · 2026-08-10
- AI Basketball Coach Robot LUMISTARCARRY Raises Over $1M on Kickstarter — 创业邦 · 2026-08-10
- Klepton: Running Android ARM64 VR APKs on Apple Vision Pro — LorenDB · 2026-08-10
- Chery's Humanoid Robot Mornine Reaches 2,000 Deliveries Across 60+ Countries — CyberRobooo · 2026-08-10