Community fine-tunes an MTP head for Bonsai 2 27B, ~1.25x inference speedup
cephaloform · x · 2026-09-18
Developer SSHCodes released an open-source MTP (multi-token prediction) head for bonsai 2 27b on Hugging Face (ProCreations/Ternary-Bonsai-2-27B-MTP). Fine-tuned from Qwen 3 27B's MTP head and adapted to bonsai, it delivers roughly a 1.25x inference speedup, usable with speculative decoding to accelerate existing model inference.
More from Infra
- Baseten launches server-side web search for open models, claiming 15% lower latency — AccBalanced · 2026-09-18
- Self-built Blackwell Colab pipeline generates a 2-hour MiniMax H3 movie for $6.96 — Interesting-Town-433 · 2026-09-18
- LaurieWired's CppCon keynote covers new memory hierarchies and how to prepare — lauriewired · 2026-09-18
- NVIDIA unpacks how CUDA's full stack powers specialized AI across finance, health and manufacturing — NVIDIA Developer · 2026-09-18
- 600 tok/s single-request on Qwen 35B with Ninfer on an RTX Pro 6000 — CharlesStross · 2026-09-18
- Cadence sees India's EDA market doubling to $7.82B by 2031 — bookwormengr · 2026-09-18