Astra scores 98%+ on 10 of 37 benchmarks, nearly saturating eval suite

dswg97 · x · 2026-09-09

The evaluator flags a key caveat: Astra is close to saturating their benchmark suite, scoring ≥98% on 10 of 37 benchmarks, while GPT-5.5 saturates only 4/37. This makes confident time-horizon estimates difficult, and suggests existing evaluations are losing discriminative power for frontier models.

Related event: Astra nears benchmark saturation, scoring 98%+ on 10 of 37(2 posts)→

Original post →

More from Models

Models channel →