Distilling an LLM into two 287M GLiNER encoders for court-decision extraction — results fall just short of the teacher
SignificantZebra5883 · reddit · 2026-10-04
A detailed engineering writeup: the author turns 5M court decisions into structured graphs by having Claude Sonnet label 700 documents in 4-sentence windows (strict JSON schemas, numbered-word positions, a "what did you miss" pass, cross-window entity IDs, 25 cleanup rules), then fine-tunes two 287M models — a GLiNER span tagger (9,699 windows, 207k phrases, fp32 after bf16 NaNs, dual LRs, weight averaging over epochs 9-14) and a multiple-choice model for entity merging, kind classification and action normalization (247k auto-generated questions). The pipeline mostly works but still underperforms the teacher LLM; the full distillation recipe, including gotchas, is documented.
More from Research
- Hyper-Connection Factory: open-source repo benchmarks 11 residual-connection variants on one LLM — ChengleiSi · 2026-10-04
- Arthur Gretton to Talk on Gradient Flows on MMD at NYC Probabilistic Modeling Workshop — ArthurGretton · 2026-10-04
- Tartan IMU Challenge Draws 131 Teams, Top 10 to Present Solutions — GhaffariMaani · 2026-10-04
- How LLMs actually work: embeddings, inference dynamics and the autoregressive loop, explained — gerardsans · 2026-10-04
- Engineer pushes back on the Platonic Representation Hypothesis hype — gerardsans · 2026-10-04
- Daimon's Tactile World Model Threads Beads at IROS by Feel, Not Just Vision — CyberRobooo · 2026-10-04