Don't trust evals: early adopter says Astra is the first model to cross his trust threshold
sandersted · x · 2026-09-04
A developer shares first-hand impressions of gpt-6 Astra: "don't trust evals" — benchmarks are often sloppy and miss what matters, and small helpful behaviors like asking clarifying questions can tank a score. Still, he argues Astra is the real deal: imperfect, but the first model to cross his personal trust threshold.
More from Models
- Databricks evals: GPT-6 Astra claims SOTA on OfficeQA Pro benchmarks, cheaper per task — downingARK · 2026-09-04
- Mathematician tests GPT-6 Astra: live Lean proof verification while writing arguments — teortaxesTex · 2026-09-04
- Tavus Launches Sparrow-2, Claiming #1 in End-of-Turn Detection and Interruption Handling — ycombinator · 2026-09-04
- ARC-AGI-3: 10x reasoning tokens cuts total cost from $48k to $26k vs medium — i_dg23 · 2026-09-04
- Researchers flag data contamination concerns in benchmark behind Astra's time-horizon score — dfrsrchtwts · 2026-09-04
- antirez: judge new models by whether they fix real blocking bugs, not three.js demos — antirez · 2026-09-04