Glean reveals enterprise model routing: 75% cost cut and open-weight shift
Latent Space · rss · 2026-08-19
With rising model costs and the popularity of open-weight models like Kimi and Qwen, model routing has become critical for enterprise AI deployment. Glean, an enterprise AI platform, details its routing strategy: dynamically selecting models for different tasks in auto mode, which makes it 4x cheaper than Claude Code ($0.45 vs $1.84 per task). Glean also introduced Waldo, an agentic search model that preprocesses data before handing off to frontier models, reducing latency by 50% and token usage by 25%. The article notes that enterprise interest in open-source models has surged in the last three months due to cost pressures.
More from Infra
- China opens world's largest AI data center targeting 1 million GPUs — teortaxesTex · 2026-08-19
- Morgan Stanley: US data centers need 68GW power by 2028 — tctjr · 2026-08-19
- Local AI apps on a MacBook at zero marginal cost via Pinokio — MLX is alive and well — cocktailpeanut · 2026-08-19
- PA Governor Enacts Nation's Strictest AI Data Center Standards via Executive Order — TinfoilTricorn · 2026-08-19
- Miles v0.1 Open Source RL Framework Launches for LLMs — ying11231 · 2026-08-19
- OpenAI Codex Lead Reveals Tokenizer Inefficiency Can Spike API Bills by 34% — 新智元 · 2026-08-19