Qwen 3.8 27B scores 17.6% on strict SlopCodeBench checks
corruptbytes · reddit · 2026-08-20
Benchmark results show Qwen 3.8 27B scoring 17.6% on strict checkpoints in the HumanLayer Opus subset, trailing DeepSeek V4 Flash and Claude Code. It struggles with strict codebase management but performs adequately on core checkpoints.
More from Models
- Analysis estimates GPT-5.6-Sol params at 1.5T to 2.5T — scaling01 · 2026-08-20
- ChatGPT Outage: OpenAI Identifies Issue, Monitoring Recovery — ns123abc · 2026-08-20
- ChatGPT Down: Login and Signup Issues Reported — ns123abc · 2026-08-20
- Models can now code entire complex software in under an hour — BLUECOW009 · 2026-08-20
- Upcoming benchmark: Local deployment comparison of Qwen3.8, Gemma4, and GPT-OSS — karminski3 · 2026-08-20
- Claude Code Faces Rough Month; Alternative Models Evaluated — Hesamation · 2026-08-20