Running Qwen3.6 35B-A3B with 131K context and vision on a 6GB RTX 2060 — full config
Szadbaverem69 · reddit · 2026-10-10
- A user runs Qwen3.6-35B-A3B (Q4KM) on an RTX 2060 6GB with 131K context and vision via llama.cpp, hitting 600 tok/s prefill / 23 tok/s decode on an empty KV cache, settling to 485/15 tok/s around 90K context.
- Key tricks: offloading 39 MoE expert layers to CPU (--n-cpu-moe 39) and running the vision projector on CPU (--no-mmproj-offload) to keep VRAM under 5.2GB; Q80 KV cache.
- Full reproducible llama-server launch command included, with sampling params and mmap+mlock settings.
More from Infra
- Bittensor GPU rental network sees 45% spend growth, 106% more rentals in monthly report — markjeffrey · 2026-10-10
- Building a pit crew for Grok Bot: frontier model plans, free models grind — alexcovo_eth · 2026-10-10
- AI boom turns into a debt boom: Oracle 5y CDS near record 261bps, implying 20.4% default odds — cyb3rops · 2026-10-10
- Fireworks AI Discloses Security Incident Involving Unauthorized Use of Internal Credentials — lqiao · 2026-10-10
- US grid adds 86GW this year while AI labs need hundreds of GW of power — FinanceYF5 · 2026-10-10
- SpaceX building industrial base to make up to 1,000 Starships a year, with Florida GigaBay 11x bigger than current Megabay — XFreeze · 2026-10-10