Local Qwen 3.8 Benchmarks: 3-9% Failures Due to Infinite Reasoning Loops
on_line187 · reddit · 2026-08-21
The author ran local benchmarks for Qwen 3.8 27B Instruct (Q80 GGUF) on dual RTX 3090s across GSM8K, MATH-500, HumanEval, and MBPP using mechanical grading (unit tests/sympy, no LLM judges).
Key Findings:
- Overall Scores: GSM8K 96.7%, MATH-500 86.4%, HumanEval 95.5%, MBPP 80.0%.
- Infinite Loop Defect: 3-9% of items triggered a non-terminating reasoning state, consuming the token budget with no output. Excluding these stalls raises MATH-500 score from 86.4% to 94%. This defect also wasted significant compute time.
- Difficulty Decline: MATH-500 scores showed a monotonic decline by difficulty level, with Geometry being the weakest subject (78.0%).
- Contamination Test: Using a continuation-similarity metric, HumanEval showed significant signs of memorization (0.387 similarity vs 0.080 baseline), while MBPP appeared relatively clean.
More from Models
- Gemini 3.7 Flash tops ARC-AGI benchmark at $0.12 per task — rakyll · 2026-08-21
- o1 excels at fixing vision-grounded bugs in rendering pipelines — teortaxesTex · 2026-08-21
- Mystery stealth model 'Ox Alph' appears on OpenRouter, China lab suspected — Neosinic · 2026-08-21
- Claude Output Occasionally Mixed with Chinese Characters — springrod · 2026-08-21
- GLM-5.3 Vision questioned as Ox Alpha source; model scale deemed non-essential — teortaxesTex · 2026-08-21
- Why people imagine Grok would have full CoT visibility — teortaxesTex · 2026-08-21