Grandma GPUs reborn: 2x Tesla P40 hits 48 tok/s on Qwen 27B via F16 cache + MTP

Jumpy-Operation-4615 · reddit · 2026-09-06

A self-described non-coder used Codex to tune a $1,100 self-built 2x Tesla P40 cluster, pushing Qwen 27B Q8 inference from 15 tok/s to up to 48 tok/s short-context decode and 440 tok/s prefill.

Key findings:

Takeaway: old datacenter GPUs have plenty of headroom; the bottleneck is configuration, not hardware.

Original post →

More from Infra

Infra channel →