Taming token costs: model routing, open LLMs, and context caching
rseroter · x · 2026-09-02
The author outlines common levers for controlling token costs: routing to cheaper models, using open LLMs, and applying context caching—typically a mix of all three. He also recommends Balaji's post on context caching in agent harnesses.
More from coding & agent
- Deploying DeepSeek-V4 on Blackwell: Fixing 3 critical SGLang bugs — shrug_hellifino · 2026-09-02
- LangChain fine-tunes Qwen for agent evals, cutting costs by 100x vs GPT-5.5 — LangChain · 2026-09-02
- New Book: Building AI Agents from Design Patterns to Production — JordiRib1 · 2026-09-02
- User praises Grok Bot for eliminating manual data entry — djcows · 2026-09-02
- Claude Code generates Raspberry Pi case in minutes via AgentCad — shekitup · 2026-09-02
- Why Prime Agent Chose RLM Trajectories Over Prompt Tuning — CShorten30 · 2026-09-02