HuatuoGPT-3: open 9B medical LLM trained with one-stage RL, no SFT
jacek2023 · reddit · 2026-09-18
FreedomIntelligence released HuatuoGPT-3-9B, a medical LLM built on Qwen3.5-9B using One-stage Policy Optimization (OnePO), which adapts models to medicine in a single RL stage without prior domain-specific SFT; teacher responses guide early training and are retired as the model improves. Training code, the OnePO-Medical-20K RL dataset, and an 8B rubric grader are all open-sourced.
More from Models
- Using an LLM as benchmark scorer fails: over-optimistic ratings diverge from human judgment — amplifiedamp · 2026-09-18
- Jev as an LLM judge flops: scores nearly everything positively, disagrees with humans — amplifiedamp · 2026-09-18
- Noam Brown: models may perform their chain of thought; alignment must be solved — infoxiao · 2026-09-18
- Jev fails as an LLM scorer on OntBench: rates almost everything positively, contradicting human and Codex ratings — amplifiedamp · 2026-09-18
- Independent eval puts new model Jev at Terra no-think level, roughly on par with Luna-xhigh — tokenbender · 2026-09-18
- Hugging Face adds zero-shot text classification pipeline to its API, no training needed — joeddav · 2026-09-18