OpenAI's 'yapping penalty' cuts Astra's HealthBench lead nearly in half after length adjustment

imjustnewatai · x · 2026-09-05

HealthBench Professional tests AI on clinician use cases: Astra scored 69.5 raw vs GPT-5.6 Sol's 64.1, but after length adjustment it's 63.4 vs 60.5 — a 5.4-point lead shrinks to 2.9, with 46% of the gap evaporating.

Key numbers: Astra's answers were 27% longer on average. OpenAI's explanation: longer answers get more chances to tick rubric boxes even when the extra text adds nothing — every 500 characters beyond 2,000 costs 1.47 points, applied uniformly to all models.

Astra still wins, just by a much smaller margin — as the author quips, "get to the point" may need its own benchmark.

Original post →

More from Models

Models channel →