Open-weight 4B models approach o3-level performance on Swedish medical exams

AccomplishedCat4770 · reddit · 2026-07-26

Small open-weight 4B models get close to o3 on Swedish medical QA

The author tested smaller open-weight LLMs on multiple-choice questions from Swedish medical licensing exams.

The author also found that unconstrained reasoning can spiral into repetitive loops, and that an early-exit thinking intervention from the S-GRPO paper helped by cutting off the reasoning trace at a preset length. A reinforcement-learning attempt to shorten reasoning traces only produced modest gains.

A curious detail: Qwen3.5-4B reasoned in English even though the prompt and exam were in Swedish, suggesting language itself was not a major barrier in this setting.

Original post →

More from Models

Models channel →