User tests reveal Qwen3.8-27B knowledge regression vs 3.6
EmPips · reddit · 2026-08-20
A user's personal benchmarks indicate that Qwen3.8-27B performs worse on factual recall than its predecessor, Qwen3.6, a finding supported by third-party offline knowledge evaluations. While the model remains strong in coding, it may be less suitable for offline scenarios relying on internal weights rather than tool-calling.
More from Models
- Stanford Framework Boosts DeepSeek Past Claude at 1/11th Cost — FuSheng_0306 · 2026-08-20
- Qwen3.8-Max tops frontend code leaderboard, beating Claude Fable 5 — Alibaba_Qwen · 2026-08-20
- Gemini 3.1 Pro Still the GOAT in Most Benchmarks Except Coding — Last_Conclusion_8984 · 2026-08-20
- Qwen3-Powered ASR Model superwhisper/s1-mini Trends on Hugging Face — superwhisper · 2026-08-20
- Qwen3.8 27B Speeds Up 3x Post-Release: Open Source Advantage — TheMoonMidas · 2026-08-20
- Mac can now run a 27B model locally that codes, reasons, and sees — TheMoonMidas · 2026-08-20