Start with Low Effort, Then Scale Up

mattpocockuk · x · 2026-07-15

The author suggests that when using a model's effort dial, you should start at a low setting and only increase it if quality drops, rather than defaulting to high effort.

The reasoning is that increasing effort essentially means injecting more tokens into the prompt. This might look great on benchmarks—spending 20% more tokens for a 2% gain is good for marketing—but it burns tokens for routine tasks like exploring code repos or modifying tests. This is especially true for more expensive models like Fable.

They also point out that using more tokens increases the risk of attention degradation, which can actually lower output quality. Those "higher effort equals better quality" charts often hide the true costs, latency, and wasted resources.

Original post →

More from Models

Models channel →