Qwen 3.8 27B vs 3.6: quality up 8% but runtime 5x longer and 4x more tokens
DerTomsn · reddit · 2026-09-04
A community benchmark on oMLX compares Qwen 3.8 27B against 3.6: quality score rises from 81.1 to 87.7 (+8%), but speed drops 16% (35→29 tok/s), runtime grows 5x (8m51s→44m39s), and output tokens jump from 18K to 78K. Better quality, but you pay heavily in time and tokens.
More from Models
- Quick Question: Does GPT-6 Include HuggingFace Access? — gordic_aleksa · 2026-09-04
- "Imagine believing in benchmarks in 2026": AI circle mocks leaderboard worship — vasuman · 2026-09-04
- The real interactivity test: learning a new language purely by talking to an LLM — akbirthko · 2026-09-04
- 753B model 'thinks', 4B model writes: latent-space handoff claimed to be 20x faster — burny_tech · 2026-09-04
- OpenAI Engineer: Anthropic's Tokenizer Change Snuck a ~30% Cost Hike into Opus 4.7 — stevenheidel · 2026-09-04
- Token pricing is meaningless now: Astra cheaper per task than Gemini 3.8 Flash — stevenheidel · 2026-09-04