LLMs excel at 1-token output — dev proposes replacing low/medium/high reasoning tiers with token counts

arkuto · reddit · 2026-09-22

A Reddit developer argues LLMs are remarkably strong when limited to a single output token (e.g., picking A-D from multiple choice), making reasoning 'overrated' for many tasks. Yet most modern models can't fully disable reasoning — the 'low' tier can burn 4,000+ thinking tokens in some models, blowing the budget for his many-small-prompts project. He proposes the industry replace vague low/medium/high/xhigh reasoning tiers with explicit min/max reasoning token counts, enforceable via banning </think> and hard truncation, for predictable costs. He also notes that tools like Jev replicate what non-reasoning LLMs already do with one-token constraints.

Original post →

More from Models

Models channel →