How do you route long-running agents across models after a cost shift?
Katleen_Cole · reddit · 2026-09-11
A Reddit user revisits agent pipeline model routing after a new model release changed capability/cost tradeoffs. The workflow spans planning, tool calls, document/image interpretation, retries and final validation, so not every step should hit the same model. He measures task success, latency, context size, retry rate and total token cost, with separate tests for KV-cache behavior on smaller models and multimodal steps, and asks what signals others use to decide which stages stay on the strongest model path.
More from coding & agent
- "Agents never say 'this is wrong, rethink the plan'" — plus Claude's song about making everyone rich — ctjlewis · 2026-09-11
- Four-step playbook for "goal-driven AI": getting better results from GPT-6 Astra agents — daniel_mac8 · 2026-09-11
- Codex /side chat may cause near-full prompt cache misses, user flags design — YouJiacheng · 2026-09-11
- CursorBench 4.0 launches; Muse Spark 1.3 matches Sol at under 40% the cost — jyangballin · 2026-09-11
- Yacine vents: 'If I hear one more person say MCP I'm going to lose my mind' — yacineMTB · 2026-09-11
- muse spark 1.3 scores strong and cheap on CursorBench 4 — infoxiao · 2026-09-11