17-Year-Old's llama.cpp Kernels Push Two $80 Tesla P100s to 60 tps on Qwen 27B

Kmic68 · reddit · 2026-09-28

A 17-year-old developer released llama.cpp kernel optimizations for Tesla P100 cards ($80 each), running a Qwen 27B at q6k with 50-60 tps decode on two GPUs: up from 7-15 tps, with 260k-context decode jumping from 2-4 to 30-35 tps and prefill from 220 to 350 tps. Fixes included mixed fp16/fp32 math to eliminate rounding errors and strict regression testing. Flags matter: -c 262144 -b 32768 -ub 1024 -np 1 or MTP acceptance collapses. Merged Sept 22 with full docs and math proofs; good cooling could add another 5-10%.

Original post →

More from Infra

Infra channel →