Fine-tuning MedGemma on one DGX Spark: 4B jumps 56%→71.6% on oncology, plus a suspected Unsloth LoRA scale bug

Few-Rough-2215 · reddit · 2026-10-06

A developer fine-tuned MedGemma 4B (LoRA) and 27B (QLoRA) for oncology on a single DGX Spark using public data, 5 datasets / 9 tasks, a frozen 2,199-item quiz and paired McNemar tests: 4B went 56.0%→71.6% in 2h39 of training, 27B 68.8%→77.9%, with the tuned 4B beating the base 27B (p=0.002). Weakest spots: exact ICD-10 codes (37.6%) and a val-to-test MCQ collapse across all models.

Key open finding: a suspected LoRA scale discrepancy under Unsloth — merging at nominal scale lost most gains (83.8% adapter agreement) while 2x alpha/r restored them (93.2%), backed by identity probes. The author released a probe script and asks for independent reproduction. Also: bf16 rounding erases 17-37% of delta elements on merge, so verbatim memorization doesn't survive. Models, quiz and report are public; research use only, not a medical device.

Original post →

More from Research

Research channel →