Fine-tuned DFlash 2 Drafter Boosts Ternary Bonsai 2 27B by 2.2x on an L4

naklitechie · reddit · 2026-10-07

A hobbyist retrained z-lab's DFlash 2 speculative-decoding drafter for PrismML's Ternary Bonsai 2 27B (30 tok/s on L4, 21 on M4 Pro), since the original drafter was trained on bf16 Qwen3.8-27B. Fine-tuned on 1.5M tokens of Bonsai 2's own greedy output:

Chat/prose is break-even; use temperature 0. Weights (safetensors + Q4KM GGUF), full llama-server commands for all three runtimes, and docs are on Hugging Face and GitHub.

Original post →

More from Infra

Infra channel →