Not Diamond proposes new benchmarks for evaluating model routing in AI agents

rohanpaul_ai · x · 2026-09-02

Not Diamond released a document detailing their benchmarking methodology for evaluating model routing in AI agents. They argue that traditional approaches like single-turn and first-turn routing are ineffective and brittle for long-horizon agent workloads due to KV cache invalidation and inability to handle mid-session complexity shifts. Not Diamond Code uses a novel routing technique that predicts future rewards and costs at each step in a cache-aware manner, taking into account session state, token counts, and task complexity. The company also shares tooling for clients to benchmark their routing solutions independently.

Original post →

More from coding & agent

coding & agent channel →