HuatuoGPT-3: open 9B medical LLM trained with one-stage RL, no SFT

jacek2023 · reddit · 2026-09-18

FreedomIntelligence released HuatuoGPT-3-9B, a medical LLM built on Qwen3.5-9B using One-stage Policy Optimization (OnePO), which adapts models to medicine in a single RL stage without prior domain-specific SFT; teacher responses guide early training and are retired as the model improves. Training code, the OnePO-Medical-20K RL dataset, and an 8B rubric grader are all open-sourced.

Original post →

More from Models

Models channel →