Self-built AtCoder benchmark on Qwen 3.8: max thinking effort mostly burns tokens for nothing
ColorsOfCosmos · reddit · 2026-10-11
A 5060Ti 16GB owner built a mini-benchmark to pick a daily driver: 14 AtCoder problems from the past 30 days (avoiding training contamination) with a LiveCodeBench-based harness, comparing three Qwen 3.8 variants — Swift 1.5 IQ2XS, Flash Next IQ3S, 27B IQ3S — across all four thinking efforts (none/low/medium/xhigh), thinking budget capped at 32k. The full sweep took 11 hours on a 7945HX + 64GB RAM + 5060Ti.
Key findings:
- Raising thinking effort buys almost nothing: all three models stay within ±2 problems of their low-effort score, with no monotonic low→med→xhi improvement (Flash peaks at med with 11, Swift at xhi with 10, 27B flat at 10).
- Higher effort mostly buys CAP-outs: 32k-cap hits rise with effort while runtime grows 33–57% — extra tokens go to problems the model would fail anyway.
- The only reliably degrading setting is thinking off: all 4 DNFs occur at none, and errors jump sharply.
Verdict: Flash IQ3S is the best speed/accuracy compromise (solved one more than 27B at medium); 27B is reliable but too slow to justify; daily driver = medium effort, escalate to xhigh only when stuck. The author acknowledges the small sample size.
More from Models
- Model Refuses to Copy a File Over Copyright, Then Changes Its Mind — burkov · 2026-10-11
- AI can decompile Morrowind to source and get it running in a browser in about a day — nitarshan · 2026-10-11
- Don't rush to switch models: benchmark on a real task and do the ROI math first — sujingshen · 2026-10-11
- GPT's 'pro-AI bias' protocol draws user backlash for downplaying controversies — GreenBird-ee · 2026-10-11
- Grok voice, one month old, beats ChatGPT voice mode that's aged two years — Scobleizer · 2026-10-11
- Gemini community rage peaks as users blast Google over endlessly withheld Argon and other models — Ok-Representative-17 · 2026-10-11