Grandma GPUs reborn: 2x Tesla P40 hits 48 tok/s on Qwen 27B via F16 cache + MTP
Jumpy-Operation-4615 · reddit · 2026-09-06
A self-described non-coder used Codex to tune a $1,100 self-built 2x Tesla P40 cluster, pushing Qwen 27B Q8 inference from 15 tok/s to up to 48 tok/s short-context decode and 440 tok/s prefill.
Key findings:
- F16 KV cache over Q8: F16 gives much better MTP speculative-decoding acceptance — the biggest single win, though it requires retuning other params.
- Runs llama.cpp/p40.cpp with locally built NCCL: tensor split 1:1, Flash Attention on, draft-mtp + ngram-simple speculative decoding with draft max 6.
- Single 220K-token context slot; a 63,900-token prefill measured 258 tok/s, and with a changed suffix the LCP prefix cache hit fkeep=1.000, recomputing only 4 tokens.
- Full production command line included, with sampling and reasoning settings.
Takeaway: old datacenter GPUs have plenty of headroom; the bottleneck is configuration, not hardware.
More from Infra
- Blacklisted Inspur Kept Buying Nvidia's Best AI Chips via US Subsidiary Aivres, NYT Finds — ShakeelHashim · 2026-09-06
- Open-source Termux scripts turn old Android phones into GPU-accelerated Linux desktops or Home Assistant hubs — tom_doerr · 2026-09-06
- Running H3 across a 3090 and unlocked 64GB CMP 170HX hits OOM in ComfyUI — JustinPooDough · 2026-09-06
- Baseten's Philip Kiely Launches Inference Engineering Book, Plus Learning Resources — kmeanskaran · 2026-09-06
- Polygres turns your Postgres into a hybrid search context layer for AI agents — Scobleizer · 2026-09-06
- TCS may invest up to $7.4 billion with TPG in a gigawatt AI campus in Hyderabad — emmanuelvivier · 2026-09-06