Start with Low Effort, Then Scale Up
mattpocockuk · x · 2026-07-15
The author suggests that when using a model's effort dial, you should start at a low setting and only increase it if quality drops, rather than defaulting to high effort.
The reasoning is that increasing effort essentially means injecting more tokens into the prompt. This might look great on benchmarks—spending 20% more tokens for a 2% gain is good for marketing—but it burns tokens for routine tasks like exploring code repos or modifying tests. This is especially true for more expensive models like Fable.
They also point out that using more tokens increases the risk of attention degradation, which can actually lower output quality. Those "higher effort equals better quality" charts often hide the true costs, latency, and wasted resources.
More from Models
- RoMa v2 image matching model unveiled in the usual black poster — ducha_aiki · 2026-09-11
- OpenAI rated Astra 'Critical' for cyber capabilities — and admits it's harder to monitor — theguywhobuilds · 2026-09-11
- TestingCatalog's Daily AI Brief adds email editions, dishing Meta Muse and GPT-Live-1 rumors — testingcatalog · 2026-09-11
- ChatGPT monthly active users top 1.06 billion in August, fourth straight record month — FinanceYF5 · 2026-09-11
- PuzzleMask: Plain-Prose Attack Bypasses All 4 Tested LLM Gatekeepers at 100% — TechNadu · 2026-09-11
- OpenAI Codex may issue another usage reset this weekend, says Codex lead resets happen — umesh_ai · 2026-09-11