NVIDIA shows how to run Gemma and Qwen locally on Jetson with Ollama and vLLM
NVIDIA Developer · youtube · 2026-07-29
NVIDIA’s Jetson AI Lab session shows how to run open-source GenAI locally on Jetson devices without cloud access.
- It walks through getting models like Gemma and Qwen running end to end on-device.
- The stream compares Ollama for rapid prototyping, vLLM for higher-throughput serving, and llama.cpp for lightweight edge deployment.
- It also covers Jetson-specific serving choices for Orin and Thor, plus practical tuning topics like quantization and speculative decoding.
- A live demo of Reachy Mini shows the stack running in real time on edge hardware.
More from Infra
- Cradle Codec: GPU-Native KV Cache Compression Explained — knowrohit07 · 2026-07-29
- Cheap local intelligence could shift AI workloads away from the cloud — PeterDiamandis · 2026-07-29
- Bull case says AMD profit could 10x as AI spend and inference demand scale — AccBalanced · 2026-07-29
- Cradle Codec compresses KV cache for Ethernet transport between GPU nodes — knowrohit07 · 2026-07-29
- Bittensor raises q from 0.61 to 0.75, easing its emission gate for mid-ranked subnets — markjeffrey · 2026-07-29
- OpenRouter’s moat comes from routing data and tooling it can refine multiple times a day — mmurph · 2026-07-29