GLM 3.8-27B quality 'absurdly superior' to 35B-A3B models, saving 22-33% tokens
JLeonsarmiento · reddit · 2026-09-12
An applied-science researcher shared a hands-on account of using GLM 3.8-27B (a Qwen-based dense model) for full research workflows, calling it 'absurdly superior' to every 3.5/3.6-35B-A3B variant (kat, Ornith/tiel, nex-2, etc.), with only Ornith coming close.
- Quality: Fully replicated 5 past projects end-to-end — workflow design, data pipelines, results analysis, report writing, and data publishing — with what the author calls a 'stupid level of attention to detail.'
- Cost: 3-4x more wall time, a real toll on daily throughput on a laptop; yet the quality gap between 3.8-27B and Z.ai's 5.3/5.3-flash is far smaller than its gap with the 35B-A3B family.
- Efficiency: At effort=medium it spends 22-33% fewer tokens with a lower RAM footprint, hitting limits and compaction less often.
The author's verdict: accept the slower speed for the quality and let the 'fat bottom Qwen' work.
More from Models
- Local Qwen Vision Model Spots Poisonous Oleander at Home, Accurately Diagnoses Skin Condition — Reasonable_Goat · 2026-09-12
- AI models keep citing a two-year-old forum thread over our official docs — OtherwiseSection7795 · 2026-09-12
- GPT-3 to Chinchilla: A Thread Walking Through the Papers That Reshaped AI — thisdudelikesAI · 2026-09-12
- DeepSeek tests voice chat, OpenAI agent accused of attacking RubyGems, Cohere seeks $3B — 快鲤鱼 · 2026-09-12
- Open-weights models keep shrinking, on-device capability replication is near — menhguin · 2026-09-12
- One-prompt galaxy collision benchmark: Opus 5 beats Fable 5.1, GPT-6 Astra finishes last — Fleischkluetensuppe · 2026-09-12