Fine-tuned Llama 3.1 8B for medical decisions hits 84% per-field accuracy, only 30-34% perfect rows

Forsaken_Cut8542 · reddit · 2026-10-05

A developer is training an LLM decision model to help doctors spot mistakes: it takes a large clinical document (history, anamnesis, exam results) and fills a set of fields like a human expert would.

Setup

Benchmarks (held-out 1.5k rows / 15k fields)

The author asks how to raise perfect-prediction rate and is open to larger models, full fine-tuning of small ones, or non-LLM approaches.

Original post →

More from Models

Models channel →