Qwen3.8-27B MTP Grafted to Unsloth Saves RAM, Requires Thinking Mode
Nyghtbynger · reddit · 2026-08-23
The author grafted the MTP (Multi-Token Prediction) mode onto the Unsloth version of Qwen3.8-27B, as the official Unsloth releases do not ship with it. Tests on Vulkan indicate this grafting method saves RAM compared to using external files. However, the model performs poorly with Thinking disabled but recovers some intelligence on cultural questions when Thinking is enabled.
More from Infra
- Mistral reportedly plans up to 1 GW of European compute capacity by 2030 — emmanuelvivier · 2026-08-23
- Contextual News Search APIs: A Deep Comparison for AI, RAG, and Research — ermanos12 · 2026-08-23
- Running Kimi K3 on 8x B300: $190 per million tokens, full cost breakdown — OtherRaisin3426 · 2026-08-23
- Nvidia AI Server Prices to Rise 15% Due to DRAM Shortage — The Decoder · 2026-08-23
- ComfyUI Node Optimization: Sparse Attention Boosts Speed by 5-20% — Zironic · 2026-08-23
- Benchmark Report: How Much Do Quants Matter on Modern Models? — KitchenAmoeba4438 · 2026-08-23