NVIDIA open-sources 3B vision-language model, claims 10x faster grounding than Qwen3-VL
Dan_Jeffries1 · x · 2026-09-30
NVIDIA released an open-source 3B vision-language model for real-time, high-quality visual grounding.
- Parallel box decoding, claimed 10x faster than Qwen3-VL
- Trained on 138M queries / 785M boxes
- Handles GUI, OCR, document layout, and dense detection
- Aimed at computer-use agents and Physical AI
Now open source.
More from coding & agent
- OpenAI DevDay live thread: 21 updates, headline release likely the always-on agent — kimmonismus · 2026-09-30
- Apodex 1.1 launches: agents can re-plan mid-task without restarting research — SucceededMind · 2026-09-30
- Neuroscience Researcher Uses Claude and GPT as Rival Research Assistants — Now Unsure Where to Publish — neuroecology · 2026-09-30
- Structure-only AI audit: Jev predicts outcomes of 2,029 real calls at AUC 0.78 for $3 — alexcovo_eth · 2026-09-30
- Replicas V3 launches: run Claude Code and Codex in cloud VMs with BYO keys — KlausCodes · 2026-09-30
- Agno adds first-class MCP publishing for agent components — pritisinghhhh · 2026-09-30