Not Diamond proposes new benchmarks for evaluating model routing in AI agents
rohanpaul_ai · x · 2026-09-02
Not Diamond released a document detailing their benchmarking methodology for evaluating model routing in AI agents. They argue that traditional approaches like single-turn and first-turn routing are ineffective and brittle for long-horizon agent workloads due to KV cache invalidation and inability to handle mid-session complexity shifts. Not Diamond Code uses a novel routing technique that predicts future rewards and costs at each step in a cache-aware manner, taking into account session state, token counts, and task complexity. The company also shares tooling for clients to benchmark their routing solutions independently.
More from coding & agent
- Speculative PTC: Overlapping tool calls with code generation for faster agents — a1zhang · 2026-09-02
- Claude Code version 2.1.257 incoming — ClaudeCodeLog · 2026-09-02
- Paying for proprietary models buys ease of work, not higher ceiling — tokenbender · 2026-09-02
- Dev claims top OSS models match proprietary ones in research ceiling — tokenbender · 2026-09-02
- SpaceXAI engineer details running a 20-agent team with GrokBot — alexcovo_eth · 2026-09-02
- Founder moves 7 Claude marketing skills to Grok Bot, builds one-person company workflow — PrajwalTomar_ · 2026-09-02