Users now expect generous inference tokens; naive limits will become table stakes
WhatTheLJW · x · 2026-09-22
saranormous argues users may not articulate it, but they get increasingly frustrated when companies hold back inference and naively restrict tokens at the cost of performance — once you've tasted scaling tokens, you just expect it. WhatTheLJW adds that once users expect it, generous inference becomes table stakes for companies.
Related event: Investor warns token-limiting AI firms will lose users(2 posts)→
More from Models
- Opus 5.5 medium beats GPT-6 Astra max 54.5% vs 53.3% at one-fifth the cost per task — haider1 · 2026-09-23
- Andriy Burkov on why Jev-style LLM calibration can't beat an LLM at real probabilities — burkov · 2026-09-23
- Claude Opus 5.5 Sweeps CursorBench at 57.8% Max, Costs 40% Less Per Task Than Opus 5 — mattyp · 2026-09-23
- Opus 5.5 series shows major step up in 3D understanding, sketch-to-simulation demo shows — _sholtodouglas · 2026-09-23
- Frontier-Model Bug Benchmark: 105 Problems Models Failed at the Start of 2026 — PawelHuryn · 2026-09-23
- Early impressions: users are warming to how Opus 5.5 writes — basedjensen · 2026-09-23