Qwen Official Guide: Tuning reasoning depth and extending context to 1M tokens
solyarisoftware · x · 2026-08-15
Qwen released a practical guide for Qwen3.8-27B, focusing on tuning the model's 'reasoning depth' and extending the native 262K context window to 1M tokens using YaRN. The guide also includes deployment references for vLLM, SGLang, TokenSpeed, and Unsloth.
More from Infra
- NVIDIA open-sources NeMo Switchyard for dynamic model routing in agent workflows — NVIDIAAI · 2026-08-15
- Vercel ranked as the world's fastest AI Gateway infrastructure — cramforce · 2026-08-15
- Qwen3.8-2.4T-A95B deployment guide: NVFP4 needs 8×B300, TP must divide 64 — Necessary_Gazelle211 · 2026-08-15
- mcpp: Auto-generate MCP servers from C++ code via reflection — karurochari · 2026-08-15
- RTX 3090 gets 35 t/s on Qwen 3.8 27B — cviperr33 · 2026-08-15
- CME to launch futures contracts tracking Nvidia H100/B100 compute costs — AccBalanced · 2026-08-15