Are IQ quants really slow on P40? User benchmarks Qwen 3.6 35B at 37-83 tok/s

otacon6531 · reddit · 2026-09-29

A r/LocalLLaMA user runs Qwen 3.6:35b IQ4 via llama.cpp on an NVIDIA P40, hitting 37-83 tok/s with MTP on; prefill starts around 600 and degrades to 300-400 on long prompts. An AI told them IQ quants are noticeably slower on P40 (citing a two-year-old thread), but they felt no slowdown moving from Q4 to IQ4 and question if that claim is outdated.

Original post →

More from Infra

Infra channel →