Astra and Grok 4.7 sit at opposite ends of accuracy and token-usage charts in CAIS eval

polynoamial · x · 2026-09-24

CAIS/Scale AI's eval site added a chart showing Astra and Grok 4.7 at opposite ends of both accuracy and token usage — one burns tokens for precision, the other does the reverse. Meta's Noam Brown quote-shared it, glad to see it on the website.

Original post →

More from Models

Models channel →