Clinical benchmarks show Astra is incremental, still trailing Anthropic's frontier
danielmckinn0n · x · 2026-09-06
The author tested OpenAI's Astra on RareBench, a clinical genetics benchmark, and concluded it's an incremental improvement over Sol rather than a new paradigm, still behind Anthropic's Fable 5.1.
Key points:
- Clinical/life-science tasks (RareBench, GeneBench-Pro, MedChemBench) remain unsaturated; the new model's intelligence is 'spiky', not a uniform leap.
- Independent testing corroborates Artificial Analysis retuning benchmarks to penalize Meta's seemingly overfit Muse Spark 1.3 Max: the Max version only marginally improves, leaving Meta behind the frontier.
- The author suggests adding RareBench to Artificial Analysis's held-out eval set.
More from Models
- Rumor: Anthropic Solved a Millennium Prize Problem, Terence Tao Responds — littmath · 2026-09-06
- Frontier AI models fix only 1 in 4 security vulnerabilities correctly, report finds — Evgenii42 · 2026-09-06
- Bold prediction: GPT-6 Luna/Terra will automate most computer work for $20/month — xhluca · 2026-09-06
- Gemini 3.8 Flash scores 73.7% on DeepSWE, up 8.2% over 3.7 Flash at same cost — burny_tech · 2026-09-06
- The holodeck may end up procedural worlds with a generative lighting and texture pass — dreamwieber · 2026-09-06
- Ethan Mollick: 'Sparks of AGI' paper deserves credit from GPT-4 to GPT-6 — emollick · 2026-09-06