FreedomIntelligence Releases HuatuoGPT-3-27B Medical LLM With OnePO
jacek2023 · reddit · 2026-09-25
FreedomIntelligence released HuatuoGPT-3-27B, a medical LLM built on Qwen3.8-27B using One-stage Policy Optimization (OnePO), which adapts models to medicine in a single RL stage without domain-specific SFT — teacher responses give temporary guidance and are retired as the model improves. They open-source the training code, the OnePO-Medical-20K RL dataset, and an 8B rubric grader, following last week's 9B release.
More from Models
- NaceAI launches Drex, a sub-6B decision model that tops the public Decision Index at 51.73 — ordax · 2026-09-25
- Model audit showdown: Astra dominates, Opus and Fable close, Grok 4.7 and GPT-6 Sol lag far behind — ivan_bezdomny · 2026-09-25
- Uncensored local model Bonzai 2 27B tops benchmarks, runs on 12GB VRAM — alexcovo_eth · 2026-09-25
- Same prompt, Opus 5.5 one-shot video generation put to a public retest with different tools — drrickio · 2026-09-25
- Users blast GPT-5.2 for rampant false crisis flags and needless helpline redirects — ryunuck · 2026-09-25
- Agent Arena Ranks 43 Models on 2M+ Real-World Agentic Tasks; Claude Fable 5.1 Tops Board — arena · 2026-09-25