Qwen 3.6 35B quant shines as a local general-purpose agent, user reports 120-140 tok/s on dual P100s

Mrinohk · reddit · 2026-10-09

A Reddit user shares hands-on experience running the Unsloth UDQ4KXL quant of Qwen 3.6 35B as a local general-purpose assistant. They find it weak at coding medium projects and prone to hallucination, with fine-tunes overly coding-focused — but surprisingly consistent at broad everyday agentic tasks: smart home control, reminders, sending grocery lists, even installing a Minecraft modpack on a remote machine over SSH.

On two used Tesla P100s it hits 120-140 t/s generation and 1000 t/s prompt processing at zero context (700 at 13k), feeling genuinely conversational.

Weaknesses: poor persona holding compared to Gemma4 and occasional factual slips. The author worries a future Qwen4 in this class will be coding-focused and slow, arguing the general AI assistant for normal people remains unsolved.

Original post →

More from Fun

Fun channel →