HuatuoGPT-3: 27B open medical LLM hits 70.1 on HealthBench, beating GPT-6 Astra
CUHKSZ · hf · 2026-10-07
The CUHK-Shenzhen team released HuatuoGPT-3, exploring RL-only domain adaptation from base models, skipping the dominant SFT+RL pipeline. They identify two failure modes of mixed-policy RL: Gradient Starvation (informative teacher tokens learned too slowly early) and Teacher-Distribution Anchoring (stale teacher outputs hindering later improvement).
Their OnePO uses Adaptive Objective Evolution to strengthen learning on informative low-probability teacher tokens and Teacher Retirement to discard teacher outputs once surpassed. With only 20K samples, OnePO hits 67.2 on HealthBench—2.7 and 7.4 points above SFT+RL and pure RL. The scaled 27B model reaches 70.1 on HealthBench and 71.4 on HealthBench Professional, surpassing GPT-6 Astra and other frontier models. Models and code are open-sourced.
More from Models
- AI Research Agents Have Radically Different Styles—Leaderboards Reduce Them to One Blind Number — ChengleiSi · 2026-10-07
- Dev: Claude wins benchmarks but Codex better at doing what I want — Kuprel · 2026-10-07
- DataCamp founder slams OpenAI Codex's 'insane' automatic resets — hugobowne · 2026-10-07
- Developer says 6.1 is underrated: handling GIS automation to 3D reconstruction well — haider1 · 2026-10-07
- Dev Complains Overnight Long-Running Tasks Keep Hitting Usage Limits — willdepue · 2026-10-07
- Gary Marcus on whether frontier LLMs can solve open math problems without symbolic harnesses — GaryMarcus · 2026-10-07