Penalizing 'wait'/'maybe' tokens lifts Qwen3.5-4B math accuracy up to +12 points

am17an · reddit · 2026-09-28

Inspired by a Meta paper on overthinking markers, a Redditor tested logit-bias penalties of -2 on 48 hedging/backtracking tokens (wait, maybe, perhaps, hmm, however, reconsider...) across quantizations of Qwen3.5-4B in llama.cpp, on 50 random MATH-500 problems.

Results:

Suppressing hesitation tokens both cuts reasoning overhead and boosts accuracy, with large gains on heavier quantizations. The post includes the full copy-pasteable logit-bias command; caveats: one model, one small test.

Original post →

More from coding & agent

coding & agent channel →