NVIDIA boosts llama.cpp throughput up to 1.9x on RTX 5090 for local agents
Scobleizer · x · 2026-09-05
NVIDIA shipped new local inference optimizations: up to 1.9x higher llama.cpp throughput on GeForce RTX 5090, 1.2x vLLM performance on RTX PRO 6000 Blackwell, and up to 1.4x on a two-system DGX Spark cluster. The optimizations, integrated with the Hugging Face ecosystem, meaningfully speed up local model agents.
More from Infra
- Will combining multiple GPUs' VRAM for local LLMs ever work out of the box? — PusheenHater · 2026-09-05
- Declarative Attention lets LLMs declare their own focus, cutting 52% of KV cache reads — eigenlaplace · 2026-09-05
- Agent outputs die when the VM sleeps: octomind's design for deliverables that survive — donk8r · 2026-09-05
- Japan to develop AI-powered satellites — AIFlow_ML · 2026-09-05
- Hybrid Compute on Mac ships with open-sourced local inference engine and PII classifier — andrewgwils · 2026-09-05
- Tesla's RIM process kills the paint shop, shrinking Cybercab factory footprint ~50% — elonmusk · 2026-09-05