Ternary Bonsai 2 27B at 1.75bpw fits an 8GB GPU, hits 93.3% accuracy in audiobook speaker-attribution test

autonoma_2042 · reddit · 2026-09-19

A detailed hands-on evaluation pits PrismML's Ternary-Bonsai-2-27B (ternary {−1,0,+1} weights at 1.75bpw, 5.95GB, dense Qwen3.8-27B base with hybrid Gated DeltaNet attention) against Gemma 26B-A4B MoE for speaker attribution in an audiobook pipeline, running on an 8GB NVIDIA T1000.

Key findings and pitfalls:

Verdict: the 5.9GB ternary dense model is viable on a single 8GB card, though long-context reasoning and speaker consistency lag. The methodology and exact flags are directly reusable for local-deployment tinkering.

Original post →

More from Infra

Infra channel →