"If These Evals Are True, We Need Far More People Working on Evals"
HarveenChadha · x · 2026-09-04
Engineer Harveen Chadha reacts to newly surfaced benchmark scores (apparently for GPT-6 Astra): if the results are real, the industry needs far more people working on model evals immediately. A short but telling take — the stronger frontier models get, the scarcer credible independent evaluation becomes.
Related event: GPT-6 Astra Launches with Benchmark Leaks: ARC-AGI-3 Hits 98.6%(41 posts)→
More from Models
- Epoch AI: GPT-6 Astra sets ECI record of 169, tops math and continual learning benchmarks — NathanpmYoung · 2026-09-04
- GPT-6 Astra system card: no-CoT time horizon up ~10x over GPT-5.6 Sol, UK AISI finds — scaling01 · 2026-09-04
- Every's Vibe Check: GPT-6 Astra Is a Big Upgrade, but Anthropic's Fable Still Has Better Product Instincts — every · 2026-09-04
- GPT-6 Astra Nukes ARC-AGI-3: Score Jumps from 8% to 63%, 98.6% with Adapter — haider1 · 2026-09-04
- Sam Altman Officially Launches GPT-6 Astra, Claiming Best-in-World Computer Use and Coding — eyishazyer · 2026-09-04
- Early user: GPT-6 Astra rebuilt Apple Park in Blender from just images — BLUECOW009 · 2026-09-04