Fine-tuned Llama 3.1 8B for medical decisions hits 84% per-field accuracy, only 30-34% perfect rows
Forsaken_Cut8542 · reddit · 2026-10-05
A developer is training an LLM decision model to help doctors spot mistakes: it takes a large clinical document (history, anamnesis, exam results) and fills a set of fields like a human expert would.
Setup
- HTML forms parsed to minimal markdown, capped at 38k chars for a 16k context window; dates/IDs stripped, personal info kept
- Base: unsloth/Meta-Llama-3.1-8B-Instruct-bnb-4bit, QLoRA rank=16, custom unsloth DDP and memory optimizations
- 27k examples, single instruction, trained and inferred on 2x Tesla V100 16G
Benchmarks (held-out 1.5k rows / 15k fields)
- 84% avg per-field accuracy; 80% on a child subset
- Only 30-34% perfect (all-fields-correct) predictions
The author asks how to raise perfect-prediction rate and is open to larger models, full fine-tuning of small ones, or non-LLM approaches.
More from Models
- Near-identical image scores, huge gaps: AI denoising must serve science, not looks — bravo_abad · 2026-10-05
- Decision Index: 70 open reproductions of Jev's Decision Model, one benchmark — bibryam · 2026-10-05
- Critic Concedes OpenAI's GPT Computer-Use Now Drives Safari Like a Human, Barely Errs — kimmonismus · 2026-10-05
- User praises 6.1 Sol /high limits but flags persistent Codex remote connectivity bugs — rudrank · 2026-10-05
- Ornith 1.5 35B-A3B hits 180 tok/s on dual 5070 Tis, 3x faster than Qwen 27B at same agent scores — Excellent-Issue-5956 · 2026-10-05
- ChatGPT to merge Chat and Work, drop model picker—users eye cancellation — Diamond_Mine0 · 2026-10-05