Qwen 3.8 27B scores 17.6% on strict SlopCodeBench checks

corruptbytes · reddit · 2026-08-20

Benchmark results show Qwen 3.8 27B scoring 17.6% on strict checkpoints in the HumanLayer Opus subset, trailing DeepSeek V4 Flash and Claude Code. It struggles with strict codebase management but performs adequately on core checkpoints.

Original post →

More from Models

Models channel →