Rippling Benchmark: Opus Edges Out, $7B OpenRouter Deal Hints at Smart Routing Value

serendip-ml · reddit · 2026-08-19

Rippling ran a benchmark testing 15 models on real payroll work with 2,100 scored runs each. Results showed untuned models scored 88.5%-89.5%, Z.ai's GLM 5.2 hit 88.7% for $621, and Anthropic's Opus 4.6 (the only prompt-tuned model) reached 91.0% for $1,453, yet still failed 9% of runs. Weeks later, a hypothetical $7B acquisition of OpenRouter by Stripe is noted. OpenRouter processes 100 trillion tokens monthly for 8M developers, earning $140M ARR. The key question raised: How much agent inference requires heavy thinking versus simple structured output? Is smart routing essential, or is a fixed role-based setup sufficient?

Original post →

More from coding & agent

coding & agent channel →