Fine-tuned DFlash 2 Drafter Boosts Ternary Bonsai 2 27B by 2.2x on an L4
naklitechie · reddit · 2026-10-07
A hobbyist retrained z-lab's DFlash 2 speculative-decoding drafter for PrismML's Ternary Bonsai 2 27B (30 tok/s on L4, 21 on M4 Pro), since the original drafter was trained on bf16 Qwen3.8-27B. Fine-tuned on 1.5M tokens of Bonsai 2's own greedy output:
- NVIDIA L4: GSM8K 2.17x, MBPP 2.17x, MATH-500 2.20x, MT-Bench 1.39x; code edits hit 3.15x with ngram-mod stacked, accuracy within 1-2 problems per set
- Mac (M4 Pro): 1.5x code completion, 1.3x chat code, 1.2x math, with an OpenAI-compatible server script via a custom Metal GEMM fork
- Browser: a WGSL port in LocalMind, 1.18x on code with identical output
Chat/prose is break-even; use temperature 0. Weights (safetensors + Q4KM GGUF), full llama-server commands for all three runtimes, and docs are on Hugging Face and GitHub.
More from Infra
- VEDA Sparse Attention cuts MiniMax H3 video gen time in half in ComfyUI with no visible quality loss — robomar_ai_art · 2026-10-07
- H200 vs multi-GPU RTX PRO 6000 Blackwell: how to pick inference hardware by budget — recentheartbroken · 2026-10-07
- Mistral release days: user reports speed slowed again with TPS around 30 — bdsqlsz · 2026-10-07
- EmbeddingGemma 2 ported to WebGPU: image-text photo search running fully in-browser — FinancialAd1961 · 2026-10-07
- Tracking LLM API model deprecations and rolling alias changes across providers — shamikhan005 · 2026-10-07
- One of the Last American Chestnut Groves to Be Destroyed for a Data Center — Promptmethus · 2026-10-07