Continued pretraining vs RAG: an accuracy and performance comparison on Qwen 3.5 4B
funJS · reddit · 2026-09-12
The author compares continued pretraining (CPT) of a Qwen 3.5 4B model against a RAG implementation on the same base model, measuring accuracy and performance to quantify the benefit of internalizing knowledge vs on-the-fly retrieval. Detailed results in the linked article — useful for anyone deciding between CPT and RAG for vertical domains.
Related event: CPT vs RAG: Testing Knowledge Internalization on Qwen 3.5 4B(3 posts)→
More from Models
- GPT-6 Astra scores 46% vs 12% for SOTA robotics VLA across 200 trials — mobav0 · 2026-09-12
- DeepSeek App Quietly Adds Four Read-Aloud Voices, Fueling New TTS Model Rumors — testingcatalog · 2026-09-12
- Watch OpenAI's Astra effortlessly drive a browser, as users ask about phone use next — infoxiao · 2026-09-12
- Researcher speculates spatial reasoning leap comes from Blender training data — yoavartzi · 2026-09-12
- Sakana's Fugu Max hits OpenRouter: multi-agent orchestration, 1M context at $2/$6 per 1M tokens — SakanaAILabs · 2026-09-12
- "Why did you nerf Astra?" Users report model degraded days after launch — Significant-Ad6970 · 2026-09-12