17-Year-Old's llama.cpp Kernels Push Two $80 Tesla P100s to 60 tps on Qwen 27B
Kmic68 · reddit · 2026-09-28
A 17-year-old developer released llama.cpp kernel optimizations for Tesla P100 cards ($80 each), running a Qwen 27B at q6k with 50-60 tps decode on two GPUs: up from 7-15 tps, with 260k-context decode jumping from 2-4 to 30-35 tps and prefill from 220 to 350 tps. Fixes included mixed fp16/fp32 math to eliminate rounding errors and strict regression testing. Flags matter: -c 262144 -b 32768 -ub 1024 -np 1 or MTP acceptance collapses. Merged Sept 22 with full docs and math proofs; good cooling could add another 5-10%.
More from Infra
- Is Agentic scores how AI-agent-ready your website is, via a single npx command — seanwbren · 2026-09-28
- TQ: calibration-free 4-bit quantization open-sourced, hits 92.4% top-1 on Qwen 27B — textclf · 2026-09-28
- Modal's free GPU glossary mini-book surfaces via Gergely Orosz — charles_irl · 2026-09-28
- Estimating 100M DAU infra for Meta Muse: 1-4GW of power, tiny $3B sandbox layer — SuB8u · 2026-09-28
- $350 Dell from 2007 beats $1500 RTX 5070 rig on agentic LLM tasks — Truth-Does-Not-Exist · 2026-09-28
- Rural town promises every household $10k if a data center gets built — JumpCrisscross · 2026-09-28