Routing Evals Are Missing, Says Elvis Saravia as OpenRouter Criticism Rages Without Evidence
omarsar0 · x · 2026-09-27
Model routing solutions like OpenRouter are being heavily criticized, but Elvis Saravia (omarsar0) argues most of the dunking lacks evidence. Quoting Alex Atallah's response (acknowledging benchmarks don't reflect real workloads and inviting users to test on their own), Saravia shares his small experiment measuring routing on real agent workloads via a custom harness — promising numbers, though he cautions they may not generalize. He sees a real opportunity: building a proper eval for routing, complex as it is (many configurations to account for), could be highly valuable for the community.
Related event: Researchers Call for Better Benchmarks as Model Routing Faces Criticism(2 posts)→
More from coding & agent
- A One-Line Prompt to Delete Dead Code From Vibe Coding, Making Agents Cheaper — gabriberton · 2026-09-27
- Quail project on building with agents: stand on battle-tested community work — charles_irl · 2026-09-27
- Using formal methods to verify agent-generated code instead of reviewing slop — sh_reya · 2026-09-27
- Apodex Open-Sources FrontierAgent Runtime Alongside 35B Mini Model — ChrisGPT · 2026-09-27
- Mid-run prompts add reproducibility and leakage-risk columns to Apodex research table — ChrisGPT · 2026-09-27
- Apodex agent turns 2024-2026 continual learning research into timeline and evidence table — ChrisGPT · 2026-09-27