Constrained decoding makes models dumber, developer argues — here's the simple math
narphorium · x · 2026-09-19
- The claim: constrained decoding (masking invalid tokens at generation time) degrades model quality.
- The mechanism: it works by stripping probability mass from invalid tokens. If the model ever assigns any probability to an invalid token, that by definition signals confusion — and simply deleting those signals pushes outputs further out of distribution.
- The author (@CompleteSkeptic) says a longer blog post on the topic is coming.
More from Models
- Users slam OpenAI's Advanced Voice Mode: non-English support feels unfinished — Angaisb_ · 2026-09-19
- Yandex releases AliceAI-80B-A3B base model trained from scratch with 262K context under Apache 2.0 — cephaloform · 2026-09-19
- Muse Spark 1.3 Gets Cheaper Contributor-Tier Optimization With Only 1-2% Benchmark Variance — alexandr_wang · 2026-09-19
- Unreleased Tencent Hunyuan 3.5 Spotted in Early Access on OnSolo, Image Quality Impresses — HeyAmit_ · 2026-09-19
- Auto-Formalizing a 76-Page Paper With Opus 5 High Would Take ~40 Days — kfountou · 2026-09-19
- Open-weight models now take 56% of production token volume, per Vercel index — cramforce · 2026-09-19