OpenAI launches GPT-6 Astra: new SOTA on HealthBench Pro at about half Sol's cost
thekaransinghal · x · 2026-09-04
OpenAI's health-focused GPT-6 Astra sets a new state of the art on HealthBench Professional, a benchmark of real clinician tasks (care consultations, medical documentation, research) deliberately built around complex, rare cases. Key numbers: Astra beats GPT-5.6 Sol's best score at its lowest reasoning effort and roughly half the cost, and on an internal eval of common consumer health questions it makes factual mistakes over 3x less often than Sol at full reasoning. OpenAI frames it as a step toward trustworthy, abundant health intelligence for patients and clinicians.
Related event: OpenAI's GPT-6 Astra Sets New Record on HealthBench Professional(3 posts)→
More from Models
- OpenAI launches GPT-6 Astra, claiming it can do anything you do on a computer — kagigz · 2026-09-04
- Databricks evals: GPT-6 Astra claims SOTA on OfficeQA Pro benchmarks, cheaper per task — downingARK · 2026-09-04
- Mathematician tests GPT-6 Astra: live Lean proof verification while writing arguments — teortaxesTex · 2026-09-04
- Tavus Launches Sparrow-2, Claiming #1 in End-of-Turn Detection and Interruption Handling — ycombinator · 2026-09-04
- ARC-AGI-3: 10x reasoning tokens cuts total cost from $48k to $26k vs medium — i_dg23 · 2026-09-04
- Researchers flag data contamination concerns in benchmark behind Astra's time-horizon score — dfrsrchtwts · 2026-09-04