Teaching local LLMs new domains: CPT + RAG experiments with full evals
funJS · reddit · 2026-09-25
A detailed writeup of experiments teaching a local model (Qwen 3.5 4B, trained with Unsloth LoRA) a new domain via continued pretraining (CPT), in four phases: (1) picking training sets that generalize to unseen questions, (2) comparing internalized knowledge (CPT) vs. RAG-injected content for reasoning, (3) combining CPT with RAG rather than treating them as competitors, and (4) a comprehensive eval strategy including SFT to force strict schema outputs for automated checks. Full findings published on the author's blog.
Related event: Hands-on: Teaching a Local 4B Model New Domain Knowledge via CPT+RAG(5 posts)→
More from coding & agent
- MCP's biggest win isn't smarter AI — it's replacing 9 point-to-point integrations with one protocol — WirelessLife · 2026-09-25
- Snorkel AI commits $3M to Open Benchmarks Grants funding agentic AI evaluations — typewriters · 2026-09-25
- Multi-Agent Coding Setups Are Mostly a Coordination Tax, Finds One-Month Experiment — Alive_Apartment6856 · 2026-09-25
- FrontierSmith lands NeurIPS spotlight: AI-synthesized coding data beats expert curation — AccBalanced · 2026-09-25
- AI Tooling Finds 8 Long-Standing Memory Leaks in libuv, 9 Fix PRs Filed — steipete · 2026-09-25
- Stack Overflow for Agents turns 3 months old: new ChatGPT plugin and privacy upgrades — pchandrasekar · 2026-09-25