Distilling an LLM into two 287M GLiNER encoders for court-decision extraction — results fall just short of the teacher

SignificantZebra5883 · reddit · 2026-10-04

A detailed engineering writeup: the author turns 5M court decisions into structured graphs by having Claude Sonnet label 700 documents in 4-sentence windows (strict JSON schemas, numbered-word positions, a "what did you miss" pass, cross-window entity IDs, 25 cleanup rules), then fine-tunes two 287M models — a GLiNER span tagger (9,699 windows, 207k phrases, fp32 after bf16 NaNs, dual LRs, weight averaging over epochs 9-14) and a multiple-choice model for entity merging, kind classification and action normalization (247k auto-generated questions). The pipeline mostly works but still underperforms the teacher LLM; the full distillation recipe, including gotchas, is documented.

Original post →

More from Research

Research channel →