Real-time VLM on Jetson AGX Thor: 10-12 scene inferences per second at ~75ms, open source
chrismatthieu · x · 2026-09-04
Developer chrismatthieu open-sourced a real-time RGB-D VLM demo running NVIDIA Jetson AGX Thor + Intel RealSense stereo depth camera:
- Everything (including the RealSense SDK) runs on GPU; the AI describes surroundings with depth in real time.
- Switching from Ollama QWEN to in-process CUDA BLIP cut latency to 70-110ms per inference—10-12 VLM scene inferences per second.
- Two backends: Ollama + qwen3.5:4b (280-350ms, richer captions) and CUDA BLIP (lowest latency); UI shows RGB, colorized depth, range stats, and optional spoken ≤8-word captions.
- Requires RealSense D400-series camera, librealsense 2.58.x built with CUDA + zero-copy, Python 3.12.
Related event: Developer Open-Sources Real-Time VLM Scene Understanding on AGX Thor(3 posts)→
More from Embodied
- Charting when humanoid robot production will surpass the human birth rate — JosephJacks_ · 2026-09-04
- Zoox robotaxis begin airport service at Las Vegas LAS — Fowe · 2026-09-04
- In-context learning hints at a "GPT moment" for robotics — chris_j_paxton · 2026-09-04
- Developer builds web-connected interactive desktop device with Claude, blown away by the model's craft — EricBuess · 2026-09-04
- Runway and NVIDIA's GWM Worlds 2 brings world models to embodied AI training — tlakomy · 2026-09-04
- Symbolic world models strike back: mu0 beats pi0.5 with 1/100 of the data — furongh · 2026-09-04