Are IQ quants really slow on P40? User benchmarks Qwen 3.6 35B at 37-83 tok/s
otacon6531 · reddit · 2026-09-29
A r/LocalLLaMA user runs Qwen 3.6:35b IQ4 via llama.cpp on an NVIDIA P40, hitting 37-83 tok/s with MTP on; prefill starts around 600 and degrades to 300-400 on long prompts. An AI told them IQ quants are noticeably slower on P40 (citing a two-year-old thread), but they felt no slowdown moving from Q4 to IQ4 and question if that claim is outdated.
More from Infra
- Qwen3.8 Flash hits 74 tok/s single-stream, 212 tok/s aggregate on one DGX Spark — open vLLM recipe — DimeRhyme · 2026-09-29
- Redditor predicts sub-$1000 device running SOTA models will spawn the next big company — Robert__Sinclair · 2026-09-29
- Only 3 of ~6,000 data center projects hit by AI buildout moratoriums: SemiAnalysis — MatthewBerman · 2026-09-29
- Cloudflare birthday week ships 8 open source updates: forge, vinext 1.0, native Rust in Workers — ritakozlov · 2026-09-29
- Google Trends' #1 US region for every query is tiny Cheyenne, Wyoming — likely bot traffic — lilyraynyc · 2026-09-29
- Sentdex has run 4B+ tokens locally on GLM 5.3 Flash — his most-used local model ever — Sentdex · 2026-09-29