NVIDIA brings local AI push to IFA 2026: 1.9x faster inference, PAIR router, RTX Spark PCs in October
NVIDIA Blog · rss · 2026-09-04
At IFA 2026, NVIDIA, Microsoft and partners announced a broad local-AI push to make agents easier to run on NVIDIA hardware.
Highlights:
- Simplified local model setup coming to Hermes Agent, OpenClaw and Perplexity Portable Computer, built on llama.cpp with NVIDIA inference optimizations
- llama.cpp delivers up to 1.9x higher throughput on RTX 5090; vLLM gains 1.2x on RTX PRO 6000 Blackwell and up to 1.4x on dual DGX Spark clusters, available via LM Studio and Ollama
- NVIDIA PAIR (Personal AI Router): free open-source tool that discovers compatible PCs on a local network and distributes AI inference across them
- RTX Spark Windows PCs arrive in October from Lenovo and Acer, with EA, Embark and Ubisoft bringing titles to the platform
August local model wave: Nemotron 3.5 Lightning (30B), Z.ai GLM-5.3-Flash, Qwen3.8 series, LTX 2.5 video, MiniMax-H3 (with FastH3 4-step distillation, 7x faster), Meta Muse Glimmer (30B coding/agentic), and DeepSeek v4 Flash (284B MoE, 13B active, runs on 2x DGX Spark) all optimized for local NVIDIA hardware.
Perplexity Portable Computer runs full workflows locally on 24GB+ RTX GPUs, keeping sensitive data on device with optional escalation to 15+ frontier cloud models.
More from Infra
- Scaling wall? Reddit argues test-time compute is the industry's new playbook — erdematar · 2026-09-05
- Tencent Hunyuan Hy4 preview: 770B total/49B active, 1M context, Apache 2.0, day-0 vLLM — aftahi_ai · 2026-09-05
- Qwen3.8 27B Quant Fits 24GB VRAM at 100k Context, Sparking Local Model Profit-Threat Debate — ChopSticksPlease · 2026-09-05
- Speechify CEO on self-built data centers, ElevenLabs leapfrog, and the $15M AI talent war — 20VC · 2026-09-05
- He Uses Local LLMs Like a 3D Printer: 12 Games, 29 Mods and Countless Tools Built Solo — Quebber · 2026-09-05
- Zhipu monetizes compute at $8-10M/MW, 5x below Anthropic and OpenAI's $40-50M/MW — zephyr_z9 · 2026-09-05