How to pick a model per step in a multi-step agent? Devs weigh routing and evals
LeviYagami · reddit · 2026-09-25
A developer running an agent with planning, tool calls, summarizing and image generation switched from one model for everything to a different model per step — better and cheaper results, but model choice feels like guesswork.
- Candidate methods: per-step custom evals, public leaderboards, task-based rankings from routed traffic (llmapi.ai), token efficiency comparisons
- Also asking whether to keep text and image/audio models behind the same API key
More from coding & agent
- My AI agent only became useful after I defined where it had to stop — Inner-Town-5609 · 2026-09-25
- Theo: $200 Claude Code plan now clearly beats Codex, weeks after trailing it — ssh4net · 2026-09-25
- Containers Aren't a Real Security Boundary: Kata Containers and Firecracker Urged for Sandboxes — andreamichi · 2026-09-25
- Reddit asks: what's the real advantage of routing smaller model jev into LLM prompts — sogo00 · 2026-09-25
- Dev ditches Amp over subscription limits, tries Claude Desktop cloud sessions — iannuttall · 2026-09-25
- One prompt, 62 seconds, under 5 cents: Xiaomi MiMo V2.6 Pro builds a full habit tracker in Claude Code — socialwithaayan · 2026-09-25