Ternary Qwen3.8-27B runs on 8GB VRAM, claims 98.2% of original smarts under $500
alexcovo_eth · x · 2026-09-19
- @0xSero reports PrismML has built a ternary quantized version of Qwen3.8-27B: 9x smaller, retaining 98.2% of the original performance, runnable on 8GB of VRAM or 16-24GB of Mac unified memory for under $500 of hardware.
- The claim is that it beats GPT-5.6-Luna High, Opus-4.6-Max and Gemini-3.1-Pro, and ties GLM-5.2 Max and Gemini-3.6-Flash.
- Caveat: separate hands-on reports say such extreme-compression models perform far worse in real use than benchmarks suggest.
More from Models
- ChatGPT Pro 20x plan back on sale after OpenAI's capacity pause — mark_k · 2026-09-19
- OrukLabs launches Resonance-2 with 31 emotion and speaking-style signals from audio — ChrisGPotts · 2026-09-19
- OpenAI launches Astra for Law; lawyer slams ZDR privacy promises as unverifiable — bgmshana · 2026-09-19
- Noam's Actual Quote: Multi-agent Credited Less Than 10% for Math Breakthrough — eliebakouch · 2026-09-19
- LLM logprobs as soft classifiers: a year-old idea finally validated — JnBrymn · 2026-09-19
- Jev benchmarks: matches production classifiers on fixed-label tasks at ~100x lower cost — ivan_bezdomny · 2026-09-19