Glean reveals enterprise model routing: 75% cost cut and open-weight shift

Latent Space · rss · 2026-08-19

With rising model costs and the popularity of open-weight models like Kimi and Qwen, model routing has become critical for enterprise AI deployment. Glean, an enterprise AI platform, details its routing strategy: dynamically selecting models for different tasks in auto mode, which makes it 4x cheaper than Claude Code ($0.45 vs $1.84 per task). Glean also introduced Waldo, an agentic search model that preprocesses data before handing off to frontier models, reducing latency by 50% and token usage by 25%. The article notes that enterprise interest in open-source models has surged in the last three months due to cost pressures.

Original post →

More from Infra

Infra channel →