AutoAWQ Author: Reproduce Bonsai 2-Class Ternary Model for ~$43k on One B300 Node in ~4 Weeks
airesearch12 · x · 2026-10-07
The author of AutoAWQ (the quantization library behind 10k models) lays out a full recipe to reproduce a PrismML Bonsai 2-class ternary model for $43k: one on-demand B300 node for about 4 weeks.
Steps: (1) Hadamard fold into every projection (block-1024 like Prism); (2) apply CAT-Q quantization (TernaryQuench reproduces it), scaling the dataset so training takes 3 days with 10-20M tokens; (3) first distillation round: KL against frozen BF16 weights, 50B tokens across every domain, 19 days; (4) on-policy distillation with 10B tokens, 6 days.
The hard parts are weight-format conversions and gathering the right datasets (mind single vs multi-turn data); the rest is now standard. He disclaims insider knowledge of Prism. Context: Underdog's Ternary was accused of being a modified, unattributed Bonsai 2.
More from Models
- Models rapidly improve at predicting experiment outcomes, may close 90% of gap by 2030 — SaxenaNayan · 2026-10-07
- Reflection Beam and Mistral Large 4 hit GLM-5.2 level, sparking distillation gap debate — Yuchenj_UW · 2026-10-07
- Mistral's new 1T-param/49B-active model beats rivals on legal benchmarks — BLUECOW009 · 2026-10-07
- Mistral CEO: Large 4 trained on our own compute, 'RL shows no sign of saturation' — sivareddyg · 2026-10-07
- Decider model gains from unmasking and more data, authors deny benchmaxxing — antoine_chaffin · 2026-10-07
- Tesla's Grok voice goes hoarse but can't hear itself — a look at AI engineering shortcuts — PTrubey · 2026-10-07