2100-Run Agent Benchmark: Grok Tops Value, GLM Nears Frontier

andreisavu · x · 2026-07-30

Rippling ran 2100 scored runs using their private agent benchmark across models from Anthropic, OpenAI, xAI, and leading open-weight players.

The results reveal that Grok is the value leader, while GLM performs remarkably close to the frontier models, suggesting that open-weight models are rapidly closing the gap with top AI labs.

Original post →

More from coding & agent

coding & agent channel →