Running Qwen3.6 27B on Tesla V100: 128K Context Config & Performance
Traditional_Bell8153 · reddit · 2026-08-08
A developer shared their configuration and performance results for running the Qwen3.6 27B model on a Tesla V100 PCIE 32Gb GPU on Reddit.
The setup uses a Q4KM quantized main model paired with a Q80 Multi-Token Prediction (MTP) draft model, tested under a massive 128K context length. The post details the specific llama.cpp parameters used, including Flash Attention, thread counts, and batch sizes, accompanied by a performance screenshot. The author is looking to compare notes with other V100 users.
More from Infra
- Google's Decade Endgame: Challenging Nvidia's AI Chip Supremacy — BorisMPower · 2026-08-08
- Local LLM Deployment: Is 128GB or 256GB System RAM Better for 200-300B Models? — Thin_Pollution8843 · 2026-08-08
- Running MiniMax Video Generation on RTX 3060 Ti Takes 30 Minutes, Sparking Discussion — Zestyclose_Mountain6 · 2026-08-08
- Optimizing Small LLMs: Why the Standard Playbook Fails Below 1.5B Params — oli266 · 2026-08-08
- Running MiniMax H3 Locally on an RTX 3060: A Hands-on Test — the_frizzy1 · 2026-08-08
- Running MiniMax H3 Fully Local on a 16GB GPU: 8-Step Video Generation & QA Lessons — Short_Regular_7191 · 2026-08-08