Omnigent's four-point playbook tames runaway LLM bills with smart model routing
matei_zaharia · x · 2026-10-07
In an Omnigent engineering post, Yuan Tang explains how debugging sessions and multi-agent runs rack up surprise LLM bills, since cost scales with tokens rather than shipped software.
The system controls spend at four points:
- Prevention: smart routing uses a lightweight LLM to classify messages as TRIVIAL/COMPLEX and blocks trivial ones (lookups, greetings, one-line fixes) from expensive models like Opus or GPT-5 via a denytrivialtoexpensivemodel policy.
- Control: progressive budgets warn before blocking, then downgrade to cheaper models instead of interrupting work.
- Visibility: a usage page breaks down spend by model, harness, and session, with custom pricing for self-hosted models.
- Safeguards: thrashing detection, scheduled-task limits, and terminal approvals.
Recommended rollout: visibility first, then routing, then budgets.
More from coding & agent
- Agentbox: open-source inbox that lets one person manage 20+ AI agents at once — Scobleizer · 2026-10-07
- Decisions API Enters Public Beta: Contextual Screen-Aware Shortcuts Make CUA Tasks ~10x Faster — stevenheidel · 2026-10-07
- Skip the basics, build the thing first — but always learn Git, advises dev — brandon_galang · 2026-10-07
- Annoyed by 24-48h usage lag, dev builds his own real-time multi-provider API usage dashboard — Atm1n9 · 2026-10-07
- How Do Teams Handle Stale Approvals When AI Agents Take Real Actions? — Goberians1 · 2026-10-07
- Dragon-IDE: Open-Source Cursor Alternative with Inter-Talking Agent Teams — DragonBallerZzz · 2026-10-07