Qwen 3.6 35B quant shines as a local general-purpose agent, user reports 120-140 tok/s on dual P100s
Mrinohk · reddit · 2026-10-09
A Reddit user shares hands-on experience running the Unsloth UDQ4KXL quant of Qwen 3.6 35B as a local general-purpose assistant. They find it weak at coding medium projects and prone to hallucination, with fine-tunes overly coding-focused — but surprisingly consistent at broad everyday agentic tasks: smart home control, reminders, sending grocery lists, even installing a Minecraft modpack on a remote machine over SSH.
On two used Tesla P100s it hits 120-140 t/s generation and 1000 t/s prompt processing at zero context (700 at 13k), feeling genuinely conversational.
Weaknesses: poor persona holding compared to Gemma4 and occasional factual slips. The author worries a future Qwen4 in this class will be coding-focused and slow, arguing the general AI assistant for normal people remains unsolved.
More from Fun
- Economist jokes about petitioning to ban AI-displaced mathematicians from taking economists' jobs — alejandroll10 · 2026-10-09
- AI Circle Jokes: Today Is 'My Data Company Is Secretly a Post-Training Lab' Day — VoidAsuka · 2026-10-09
- spawn games can now be played right inside X posts — majidmanzarpour · 2026-10-09
- Full boot-up demo of Chromatic shown ahead of Dev Day — pvncher · 2026-10-09
- AI models play a js13k 2D platformer: Qwen3.8-Flash-Next and Clef clear 8 of 14 levels — Felladrin · 2026-10-09
- A map made by punching holes in a notebook — playable in-browser with Three.js — enjalot · 2026-10-09