Optimal Settings for Qwen3.6-27B Coding
starkruzr · reddit · 2026-07-18
The poster is discussing the most effective combinations of harnesses, server parameters, and system prompts when using **Qwen3.6-27B for coding tasks**. They first shared their llama-server presets used on a machine with 2×16GB GPUs, detailing various pairings of quantization, MTP, KV cache, and context lengths. They summarized a few key findings: - **q8_0 KV** provides roughly 1.5× more context than f16 KV with minimal quality loss. - **MTP (draft spec-decode)** speeds up generation but sacrifices some context length. - A 32GB VRAM setup can essentially only load a single 27B model at a time, and switching models incurs a few seconds of reload overhead. The core question is: given these hardware constraints, how are others configuring their harnesses, system prompts, and inference settings to make this model act as a reliable coder?
More from coding & agent
- Codex turns out 123 screensavers in one playful batch — intellectronica · 2026-07-21
- Grok Build adds `grok doctor`, resumable sessions and remote image paste — mark_k · 2026-07-21
- Autoresearch proposes packaging ML runs as studies with questions, analysis, and code diffs — morgymcg · 2026-07-21
- CHAP defines approvals, handoffs, and audit logs for human-agent workflows — DeliveryTechnical199 · 2026-07-21
- The author says Codex reached 20x and is now debugging spec decoding on a hybrid parallel setup — TheZachMueller · 2026-07-21
- Axcess adds an MCP connector for WCAG accessibility checks that scanners miss — modelcontextprotocol · 2026-07-21