Match Claude model and effort level to the task to cut token usage 2-3x
ifioknkem · x · 2026-10-02
Most people run every task on one model with High or Max effort, burning limits 2-3x faster than needed. The author's setup: Haiku 4.5 for simple tasks and formatting, Sonnet 5 for everyday work and writing, Opus 5 for complex reasoning and strategy, Fable 5.1 only for the hardest jobs. Effort levels: Low for quick summaries, Medium for daily tasks, High for complex analysis, Max only as a last resort. The core rule: deliberately match both model and effort to task difficulty.
Related event: Three Tricks to Cut Claude Token Usage in Half(2 posts)→
More from Models
- JevBench to add evals for LLM routing, RAG retrieval, and moderation use cases — airesearch12 · 2026-10-02
- Google's Argon battle-tested by 200k+ Googlers daily, not benchmaxxed — Zergylord · 2026-10-02
- JEV claims to be the first System One model hosted in the EU — juanviera23 · 2026-10-02
- Gemini knew a user's mom's name unprompted, then gave three conflicting explanations — Exact_Firefighter864 · 2026-10-02
- GPT-6.1 Sol nearly matches Astra at a quarter of the price on RareBench — danielmckinn0n · 2026-10-02
- New results from PostTrainBench v1.2 are in — mariofilhoml · 2026-10-02