HuatuoGPT-3: 27B open medical LLM hits 70.1 on HealthBench, beating GPT-6 Astra

CUHKSZ · hf · 2026-10-07

The CUHK-Shenzhen team released HuatuoGPT-3, exploring RL-only domain adaptation from base models, skipping the dominant SFT+RL pipeline. They identify two failure modes of mixed-policy RL: Gradient Starvation (informative teacher tokens learned too slowly early) and Teacher-Distribution Anchoring (stale teacher outputs hindering later improvement).

Their OnePO uses Adaptive Objective Evolution to strengthen learning on informative low-probability teacher tokens and Teacher Retirement to discard teacher outputs once surpassed. With only 20K samples, OnePO hits 67.2 on HealthBench—2.7 and 7.4 points above SFT+RL and pure RL. The scaled 27B model reaches 70.1 on HealthBench and 71.4 on HealthBench Professional, surpassing GPT-6 Astra and other frontier models. Models and code are open-sourced.

Original post →

More from Models

Models channel →