"If These Evals Are True, We Need Far More People Working on Evals"

HarveenChadha · x · 2026-09-04

Engineer Harveen Chadha reacts to newly surfaced benchmark scores (apparently for GPT-6 Astra): if the results are real, the industry needs far more people working on model evals immediately. A short but telling take — the stronger frontier models get, the scarcer credible independent evaluation becomes.

Related event: GPT-6 Astra Launches with Benchmark Leaks: ARC-AGI-3 Hits 98.6%(41 posts)→

Original post →

More from Models

Models channel →