NVIDIA brings local AI push to IFA 2026: 1.9x faster inference, PAIR router, RTX Spark PCs in October

NVIDIA Blog · rss · 2026-09-04

At IFA 2026, NVIDIA, Microsoft and partners announced a broad local-AI push to make agents easier to run on NVIDIA hardware.

Highlights:

August local model wave: Nemotron 3.5 Lightning (30B), Z.ai GLM-5.3-Flash, Qwen3.8 series, LTX 2.5 video, MiniMax-H3 (with FastH3 4-step distillation, 7x faster), Meta Muse Glimmer (30B coding/agentic), and DeepSeek v4 Flash (284B MoE, 13B active, runs on 2x DGX Spark) all optimized for local NVIDIA hardware.

Perplexity Portable Computer runs full workflows locally on 24GB+ RTX GPUs, keeping sensitive data on device with optional escalation to 15+ frontier cloud models.

Original post →

More from Infra

Infra channel →