Building evals for a domain-tuned 4B local LLM: JSON schema plus exact, partial, and semantic matching
funJS · reddit · 2026-09-24
A practitioner shares how they evaluated a Qwen 3.5 4B model after continued pretraining on a new domain (subway route planning):
- Outputs constrained to a strict JSON schema of travel legs for n-transfer journeys
- A post-CPT Unsloth SFT pass (alpaca format) improved reliable schema compliance
- Evals measure exact matches plus partial and semantic matches
They report the approach works well, with a write-up linked in the post.
Related event: Practical Guide: Designing Evals for Continually Pretrained Small Models(4 posts)→
More from coding & agent
- Setting Astra's Thinking to 'Extra High' Makes It Over-Engineer Unit Tests — astralmatrix · 2026-09-24
- Free O'Reilly Book Offers a Pragmatic Framework for Scaling AI in Engineering Teams — blaizedsouza · 2026-09-24
- Zilliz CTO: agents make the enterprise data layer impossible to ignore — No_Engineer_1224 · 2026-09-24
- Running Android emulator + Chrome with 60fps streaming in a $0.072/hr cloud VM for always-on agents — cem2ran · 2026-09-24
- 10 agent reruns reached the right neighborhood, none reproduced the key observation — rohanpaul_ai · 2026-09-24
- Using the Jev model for offensive security: OPSC ranking and sensitive file detection in Mythic — dyn___ · 2026-09-24