Qwen3.6 27B Tested on RTX 3090: 60 tok/s Achieved
Top_Outlandishness78 · reddit · 2026-07-15
A post tested Qwen3.6 27B on an RTX 3090.
The author reports achieving about 60 tok/s, with two slots showing good quality and stable tool calls; no serious coding yet.
They also note: each slot is allocated 100k KV cache, totaling about 21GB VRAM usage.
More from Infra
- Two-hour workshop covers open models, benchmark cheating, reward hacking and quantization — danielhanchen · 2026-07-21
- Nativ brings local AI model running to Mac with a desktop app and localhost API — Simon Willison · 2026-07-21
- Octen says agent search now runs at 62ms P50 with only a 6ms P90 gap — aakashgupta · 2026-07-21
- Zhipu acquires a compiler-team spinout to optimize AI inference on domestic chips — zephyr_z9 · 2026-07-21
- Open reproduction of Meta’s REWIRE data pipeline cuts the cost to about $11 — vanstriendaniel · 2026-07-21
- NeurIPS 2026 workshop will focus on on-device intelligence and local execution — YiMaTweets · 2026-07-21