Open-weight 4B models approach o3-level performance on Swedish medical exams
AccomplishedCat4770 · reddit · 2026-07-26
Small open-weight 4B models get close to o3 on Swedish medical QA
The author tested smaller open-weight LLMs on multiple-choice questions from Swedish medical licensing exams.
- On MedQA-SWE, GPT-4 reached 84% in 2024 and o3 reached 88% in 2025 on a smaller overlapping set.
- With SFT on earlier exam years, MedGemma-1.5-4B reached a passing score of 60% on the final-year exam.
- Newer open-weight models did much better out of the box: Gemma4-E4B and Qwen3.5-4B hit 77% without post-training.
- With reasoning enabled, Qwen3.5-4B reached 87% accuracy.
The author also found that unconstrained reasoning can spiral into repetitive loops, and that an early-exit thinking intervention from the S-GRPO paper helped by cutting off the reasoning trace at a preset length. A reinforcement-learning attempt to shorten reasoning traces only produced modest gains.
A curious detail: Qwen3.5-4B reasoned in English even though the prompt and exam were in Swedish, suggesting language itself was not a major barrier in this setting.
More from Models
- Grok 4.5 tests suggest large tasks work better when broken into low-poly chunks — Daniel_Farinax · 2026-07-26
- A blunt defense of open-weight models says blocking others from releasing them is “evil” — rdesh26 · 2026-07-26
- Kimi K3 is expected to go open-weight tomorrow, boosting the open-source camp — Hot_Example_4456 · 2026-07-26
- Claude usage screenshot shows Max-plan limits and $1,775.98 in credits consumed — letandrewcook · 2026-07-26
- Claude Opus 5 scores 86.3% on WeirdML v2 and still averages 7,000-plus tokens — xeophon · 2026-07-26
- ChatGPT’s math output looks like a serious paper after a two-hour prompt — airkatakana · 2026-07-26