Ternary LLMs making a comeback? Multiple 1.58-bit models released, some beat full-precision baselines

Individual-Dot5488 · reddit · 2026-08-16

Reddit users discuss the resurgence of ternary (1.58-bit) LLMs. In the past month, several labs released ternary models: PrismML's 27B (benchmarks inconsistent), Deepgrove's maple-20b-a1b (runs at 100 tok/s on iPhone), and Doses AI's medical pestle-27b-ternary (outperforms medgemma-27b on medical benchmarks at 8x smaller size). Common issue: long-horizon agentic coding, but users attribute this to lack of RL optimization in new labs, which they plan to address. Hope for a competitive ternary MoE model against Qwen3.8.

Original post →

More from Models

Models channel →