Running 35B Models at 80 tok/s on a P40
otacon6531 · reddit · 2026-07-16
The author shares benchmarks of running Qwen3.6:35B UD Q4KM on a single NVIDIA P40 24GB, reportedly achieving a burst speed of 80 tok/s while supporting a 100k context.
Key configurations include:
- Using Unsloth's model weights and TheTom's TurboQuant build of llama.cpp.
- Enabling TurboQuant without downgrading to Turbo2, in order to preserve agent quality.
- Disabling reasoning to prevent the model from getting stuck in a thinking loop.
- Skipping the vision component to save roughly 300MB of VRAM, thereby freeing up space for the full context window.
The author's goal is to maximize the performance of this older GPU without degrading the model's capabilities.
More from Infra
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- DeepSeek-V4-Flash tops out at 770 tok/s on one B300 in a vLLM batch test — Moreh · 2026-07-22
- NVIDIA starts shipping 102.4 Tbps Spectrum-6 switches for Vera Rubin AI factories — nvidia · 2026-07-22
- Apple publishes SOC 3 audit reports for Private Cloud Compute — throwfaraway4 · 2026-07-22
- Reddit GPU renters say existing platforms only give you two of three: code, recovery, fair billing — legendpizzasenpai · 2026-07-22
- The Sandboxing Manifesto: Secure Execution Environments for Agents — spirosoik · 2026-07-22