Fireworks launches Nexus: smart routing to open models cuts coding AI bills in half with 95%+ cache hits
sophiamyang · x · 2026-10-07
Fireworks AI launched Nexus, an enterprise platform for optimizing AI coding spend:
- Zero-friction setup: one FireConnect CLI command plugs into Claude Code, Codex, OpenCode and more, with SSO support, no proxy, fully reversible.
- FireRouter smart routing: each request is scored and routed to the optimal model — easy tasks go to open models, hard ones pass through to closed providers on your own key. Claims 95%+ cache hit rate at half price per token, 54% overall spend saved, and 33% savings per merged PR.
- Controls: per-user budgets, spend tracking by day/user/model/API key, 100+ TPS Fast/Priority tiers, and ROI metrics like blended cost per token.
- Cites a Faros AI production-repo case where frontier open models handle 80%+ of engineering tasks (GLM-5.2 beating Opus 4).
More from coding & agent
- Agentbox: open-source inbox that lets one person manage 20+ AI agents at once — Scobleizer · 2026-10-07
- Decisions API Enters Public Beta: Contextual Screen-Aware Shortcuts Make CUA Tasks ~10x Faster — stevenheidel · 2026-10-07
- Skip the basics, build the thing first — but always learn Git, advises dev — brandon_galang · 2026-10-07
- Annoyed by 24-48h usage lag, dev builds his own real-time multi-provider API usage dashboard — Atm1n9 · 2026-10-07
- How Do Teams Handle Stale Approvals When AI Agents Take Real Actions? — Goberians1 · 2026-10-07
- Dragon-IDE: Open-Source Cursor Alternative with Inter-Talking Agent Teams — DragonBallerZzz · 2026-10-07