Building evals for a domain-tuned 4B local LLM: JSON schema plus exact, partial, and semantic matching

funJS · reddit · 2026-09-24

A practitioner shares how they evaluated a Qwen 3.5 4B model after continued pretraining on a new domain (subway route planning):

They report the approach works well, with a write-up linked in the post.

Related event: Practical Guide: Designing Evals for Continually Pretrained Small Models(4 posts)→

Original post →

More from coding & agent

coding & agent channel →