Why yes/no answers are fast for LLMs: output tokens dominate latency
tinyfool · x · 2026-09-21
tinyfool explains LLM latency: models emit tokens one by one, so the more output, the slower. Tools like Jev are fast because they emit almost nothing — the upstream reasoning still costs time, but the answer is a tiny number of tokens. Likewise, asking an LLM a question slowly yields a long answer, while forcing a yes/no response is much quicker.
More from Models
- Chinese open models reportedly just 4.4 months behind US frontier models — emmanuelvivier · 2026-09-21
- StepFun unveils Step 5 Preview: 600B-param MoE with 1M-token context — emmanuelvivier · 2026-09-21
- MiMo near-SOTA on DeepSWE with just ~$2.6M RL run: will data cost more than training? — my_cat_can_code · 2026-09-21
- humansand's Persimmon model learns to share info gradually like humans, with Trickle Test — niloofar_mire · 2026-09-21
- ChatGPT reportedly removes free-tier chat limits, offering unlimited text chats — nikola_mr64990 · 2026-09-21
- Users say top-tier Astra is too costly, hope GPT-6 fixes token economics — CtrlAltDwayne · 2026-09-21