Llama.cpp PR Moves Sampling to GPU, Boosting Inference Speed by 12% on RTX 5090

otacon6531 · reddit · 2026-08-04

A recent Llama.cpp PR moves the sampling process from the CPU to the GPU, significantly improving inference speed for users with MTP (Multi-Token Prediction) enabled.

Benchmarks show a roughly 12% speed increase running Qwen3.6:35b on an RTX 5090, and a 4% gain on an older Tesla P40. The author notes that while the P40 improvement is limited by memory bandwidth compared to high-end GPUs, it still represents one of the most substantial local inference optimizations recently.

Original post →

More from Infra

Infra channel →