FreedomIntelligence Releases HuatuoGPT-3-27B Medical LLM With OnePO

jacek2023 · reddit · 2026-09-25

FreedomIntelligence released HuatuoGPT-3-27B, a medical LLM built on Qwen3.8-27B using One-stage Policy Optimization (OnePO), which adapts models to medicine in a single RL stage without domain-specific SFT — teacher responses give temporary guidance and are retired as the model improves. They open-source the training code, the OnePO-Medical-20K RL dataset, and an 8B rubric grader, following last week's 9B release.

Original post →

More from Models

Models channel →