Self-built AtCoder benchmark on Qwen 3.8: max thinking effort mostly burns tokens for nothing

ColorsOfCosmos · reddit · 2026-10-11

A 5060Ti 16GB owner built a mini-benchmark to pick a daily driver: 14 AtCoder problems from the past 30 days (avoiding training contamination) with a LiveCodeBench-based harness, comparing three Qwen 3.8 variants — Swift 1.5 IQ2XS, Flash Next IQ3S, 27B IQ3S — across all four thinking efforts (none/low/medium/xhigh), thinking budget capped at 32k. The full sweep took 11 hours on a 7945HX + 64GB RAM + 5060Ti.

Key findings:

Verdict: Flash IQ3S is the best speed/accuracy compromise (solved one more than 27B at medium); 27B is reliable but too slow to justify; daily driver = medium effort, escalate to xhigh only when stuck. The author acknowledges the small sample size.

Original post →

More from Models

Models channel →