Start with Low Effort, Then Scale Up
mattpocockuk · x · 2026-07-15
The author suggests that when using a model's effort dial, you should start at a low setting and only increase it if quality drops, rather than defaulting to high effort.
The reasoning is that increasing effort essentially means injecting more tokens into the prompt. This might look great on benchmarks—spending 20% more tokens for a 2% gain is good for marketing—but it burns tokens for routine tasks like exploring code repos or modifying tests. This is especially true for more expensive models like Fable.
They also point out that using more tokens increases the risk of attention degradation, which can actually lower output quality. Those "higher effort equals better quality" charts often hide the true costs, latency, and wasted resources.
More from Models
- China’s AI arms race is increasingly defined by chips, data centers, and open models — BenBajarin · 2026-07-22
- Sam Altman is headed to Washington to brief Congress on OpenAI’s GPT-6 line — inductionheads · 2026-07-22
- Benchmark chart pits GPT-5.6 Luna, Grok 4.5 and Gemini 3.6 Flash on price and scores — iruletheworldmo · 2026-07-22
- Claim says Kimi was distilled from Fable, sparking a model-attribution jab — cephaloform · 2026-07-22
- Gemini 3.6 Flash is now available in Antigravity and chat — MartianOnJupiter · 2026-07-22
- Gary Marcus says LLM math skills are like knowing only a car’s engine size — GaryMarcus · 2026-07-22