Local Qwen 3.8 Benchmarks: 3-9% Failures Due to Infinite Reasoning Loops

on_line187 · reddit · 2026-08-21

The author ran local benchmarks for Qwen 3.8 27B Instruct (Q80 GGUF) on dual RTX 3090s across GSM8K, MATH-500, HumanEval, and MBPP using mechanical grading (unit tests/sympy, no LLM judges).

Key Findings:

Original post →

More from Models

Models channel →