Continued Pretraining vs RAG: Hands-on Comparison on a Qwen 4B Model
funJS · reddit · 2026-09-12
The author ran a hands-on experiment comparing two ways to feed a model domain knowledge: continued pretraining (CPT) on Qwen 3.5 4B versus a RAG implementation over the base model, measuring accuracy and performance to quantify internalizing knowledge vs retrieving on-the-fly. The write-up covers methodology and findings, offering a useful reference for developers deciding between the two approaches on private data: CPT is costlier but needs no retrieval pipeline, while RAG updates knowledge more flexibly.
Related event: CPT vs RAG: Testing Knowledge Internalization on Qwen 3.5 4B(4 posts)→
More from Research
- Embeddings decode real-time video of what a mouse saw from brain activity — flngr · 2026-09-12
- Harnesses still matter: why models won't figure out their own scaffolding — omarsar0 · 2026-09-12
- LeVJEPA: video pretraining at 5.6-20.8x less compute, matching V-JEPA 2 — Cohere · 2026-09-12
- Continued pretraining vs RAG: an accuracy and performance comparison on Qwen 3.5 4B — funJS · 2026-09-12
- Researcher speculates spatial reasoning leap comes from Blender training data — yoavartzi · 2026-09-12
- tszzl pushes back on Fermi paper: 50-OOM lognormal abiogenesis prior under-justified — tszzl · 2026-09-12