Uber Burned 2026 AI Budget in 4 Months, Then Cut Token Costs via 4 Optimizations
femke_plantinga · x · 2026-08-11
Uber exhausted its entire 2026 AI budget in just 4 months. Surprisingly, instead of restricting access, they quadrupled the number of engineers using AI tools daily while simultaneously driving down the cost per token.
They implemented four key optimizations:
- Prompt Caching & Reuse: Cache stable context and only pay for new content, avoiding the high cost of repeatedly sending the same context.
- Right-sized Defaults: Tune default models and context windows to the actual job, avoiding the habit of defaulting to expensive frontier models for everything.
- Real-time Cost Visibility: Give every engineer live visibility into their token spend, changing behavior through transparency.
- Open-weight Models & Task Routing: Test open-weight models and automatically route tasks to the cheapest capable model.
More from coding & agent
- Open-source project adds 239 design skills and 88 commands to Claude Code and Gemini CLI — tom_doerr · 2026-08-11
- Developer Test: Automating Frontend Testing and DNS Config with Codex — iannuttall · 2026-08-11
- Dev Builds 3D Arena Battler Game Using Cursor + Grok — chongdashu · 2026-08-11
- Ditch Keyboards: Autonomous Launches $149 Physical Console to Harness AI Coding Agents — dee_hw · 2026-08-11
- Git Pre-commit Hook Quizzes Devs on PR Diffs, Blocks Merges Under 85% — rudrank · 2026-08-11
- ColaMD v1.8.1 Released: Markdown Editor with Real-Time AI Agent Sync — oran_ge · 2026-08-11