OpenAI's 'yapping penalty' cuts Astra's HealthBench lead nearly in half after length adjustment
imjustnewatai · x · 2026-09-05
HealthBench Professional tests AI on clinician use cases: Astra scored 69.5 raw vs GPT-5.6 Sol's 64.1, but after length adjustment it's 63.4 vs 60.5 — a 5.4-point lead shrinks to 2.9, with 46% of the gap evaporating.
Key numbers: Astra's answers were 27% longer on average. OpenAI's explanation: longer answers get more chances to tick rubric boxes even when the extra text adds nothing — every 500 characters beyond 2,000 costs 1.47 points, applied uniformly to all models.
Astra still wins, just by a much smaller margin — as the author quips, "get to the point" may need its own benchmark.
More from Models
- Reddit user on $200 Pro plan: Astra burns weekly usage about 2x faster than Sol — TheVibrantYonder · 2026-09-05
- EnactraAI corrects cost chart, says GPT-6 Astra holds a cost edge too — Lianhuiq · 2026-09-05
- ASTRA advice: just use Medium, go High only if it fails, skip Ultra — keyanzhang · 2026-09-05
- Scale AI's Muse Spark 1.3 Max Nears Frontier Performance on Efficiency Frontier — jffwng · 2026-09-05
- GPT-6 Astra Computer Use: screenshot-driven control with mid-task steering — xiaohu · 2026-09-05
- Meta offers 95% discount on new AI model in exchange for watching you use it — technextpreneur · 2026-09-05