CPT a 9B model on 2B legal tokens: how do you restore instruct and thinking?

SignificantZebra5883 · reddit · 2026-10-07

A developer planning continued pretraining of Qwen3.5-9B on a 2B-token legal corpus asks how to turn the base model back into an instruct + thinking assistant: whether distillation from the instruct model is the only route, which European-language datasets work, whether a 9B model can be agentic in a custom harness (outputting Python with built-in vectorsearchlaws() and graphsearch() functions), and how to synthesize training data from raw court decisions and statutes for harness-specific agentic RL.

Original post →

More from Research

Research channel →