Best Practices for Continual Pretraining (CPT): A Curated Resource List
liuzhuang1234 · x · 2026-08-09
The author shared deep reading notes on Continual Pretraining (CPT) to outline best practices for building domain-specialized open LLMs.
The post curates highly valuable practical resources and model reports, including:
- CPT best practices by Databricks
- Reports for Composer 2 and Code Llama
- DeepSeek-Coder-v2 and Qwen-2.5-Coder
- DeepSeekMath and Nemotron-MIND
More from Research
- Has LLM Eaten Causal Inference? Zero Causality Workshops at NeurIPS — Beautiful_Baker_2233 · 2026-08-09
- Nature Neuroscience: Compositionality Not Uniquely Human, LLMs Achieve It via Scale — SussilloDavid · 2026-08-09
- Synthetic Data Lacks a Moat; Human Expertise Remains the Core Value — himanshustwts · 2026-08-09
- NeurIPS OpenReview Timeline Updates: Does 'Modified' Indicate Rating Changes? — Responsible-Read-138 · 2026-08-09
- Goodfire CTO on Concept Manifold Geometry & a $1000/Month ML Research Agent — The Cognitive Revolution · 2026-08-09
- Goodfire CTO on Concept Manifold Geometry and a $1000/Month ML Research Agent — The Cognitive Revolution · 2026-08-09