Llama.cpp PR Moves Sampling to GPU, Boosting Inference Speed by 12% on RTX 5090
otacon6531 · reddit · 2026-08-04
A recent Llama.cpp PR moves the sampling process from the CPU to the GPU, significantly improving inference speed for users with MTP (Multi-Token Prediction) enabled.
Benchmarks show a roughly 12% speed increase running Qwen3.6:35b on an RTX 5090, and a 4% gain on an older Tesla P40. The author notes that while the P40 improvement is limited by memory bandwidth compared to high-end GPUs, it still represents one of the most substantial local inference optimizations recently.
More from Infra
- Runware launches modular Sonic Inference Pod for portable data centers, promising lower cost and faster deployment — RebeccaBellan · 2026-08-04
- Zilla 2.0 Natively Integrates MCP: Bridging REST and Kafka Without Wrappers — jkriket · 2026-08-04
- Defending Data Centres: A Progressive Case for AI Infrastructure Economics — david_stillwell · 2026-08-04
- Run a Local LLM on Your MacBook with Two Commands for Free — tomcrawshaw01 · 2026-08-04
- Dev Explores Turnkey AIO Cooling Solutions for AMD MI100 — psychoOC · 2026-08-04
- AI Chip Startup DeepX Secures Fresh Funding at 4x Valuation — pstAsiatech · 2026-08-04