Ternary LLMs making a comeback? Multiple 1.58-bit models released, some beat full-precision baselines
Individual-Dot5488 · reddit · 2026-08-16
Reddit users discuss the resurgence of ternary (1.58-bit) LLMs. In the past month, several labs released ternary models: PrismML's 27B (benchmarks inconsistent), Deepgrove's maple-20b-a1b (runs at 100 tok/s on iPhone), and Doses AI's medical pestle-27b-ternary (outperforms medgemma-27b on medical benchmarks at 8x smaller size). Common issue: long-horizon agentic coding, but users attribute this to lack of RL optimization in new labs, which they plan to address. Hope for a competitive ternary MoE model against Qwen3.8.
More from Models
- Gemini 3.7 Flash sees no performance gains in Tibetan/Sanskrit/Chinese — SebastianNehrd2 · 2026-08-16
- Using /goal command spikes token usage 13x for minimal score gain — zainhas · 2026-08-16
- Sonnet uses 2x tokens of Sol/Kimi K3 for no performance gain — zainhas · 2026-08-16
- Grok's 'Auto' mode praised for balancing speed and reasoning — mark_k · 2026-08-16
- Community debates if Qwen 3.8 27B outperforms the larger 3.5 122B — MackThax · 2026-08-16
- Claude Sonnet 5 numbers crush Sonnet 4.5 in comparison — zainhas · 2026-08-16