User tests reveal Qwen3.8-27B knowledge regression vs 3.6

EmPips · reddit · 2026-08-20

A user's personal benchmarks indicate that Qwen3.8-27B performs worse on factual recall than its predecessor, Qwen3.6, a finding supported by third-party offline knowledge evaluations. While the model remains strong in coding, it may be less suitable for offline scenarios relying on internal weights rather than tool-calling.

Original post →

More from Models

Models channel →