Teaching local LLMs new domains: CPT + RAG experiments with full evals

funJS · reddit · 2026-09-25

A detailed writeup of experiments teaching a local model (Qwen 3.5 4B, trained with Unsloth LoRA) a new domain via continued pretraining (CPT), in four phases: (1) picking training sets that generalize to unseen questions, (2) comparing internalized knowledge (CPT) vs. RAG-injected content for reasoning, (3) combining CPT with RAG rather than treating them as competitors, and (4) a comprehensive eval strategy including SFT to force strict schema outputs for automated checks. Full findings published on the author's blog.

Related event: Hands-on: Teaching a Local 4B Model New Domain Knowledge via CPT+RAG(5 posts)→

Original post →

More from coding & agent

coding & agent channel →