CPT a 9B model on 2B legal tokens: how do you restore instruct and thinking?
SignificantZebra5883 · reddit · 2026-10-07
A developer planning continued pretraining of Qwen3.5-9B on a 2B-token legal corpus asks how to turn the base model back into an instruct + thinking assistant: whether distillation from the instruct model is the only route, which European-language datasets work, whether a 9B model can be agentic in a custom harness (outputting Python with built-in vectorsearchlaws() and graphsearch() functions), and how to synthesize training data from raw court decisions and statutes for harness-specific agentic RL.
More from Research
- MEMOIR-VLM hits 91.3% balanced accuracy on dementia classification with missing brain-scan inputs — PTenigma · 2026-10-07
- Claude paper's prompter was an Anthropic employee, authors were just digesters, says user — analisereal · 2026-10-07
- Claude math result row: prompter was an Anthropic employee, authors say they were the "digesters" — analisereal · 2026-10-07
- ChipForge NPU Challenge Goes Live: Crowdsource AI Accelerator Designs Headed to Real Silicon — bittingthembits · 2026-10-07
- SIGGRAPH Asia 2026 Lineup: Disney's Markus Gross, Huawei's Ran Huang, Oxford's Andrea Vedaldi — AlexTensor · 2026-10-07
- CoreWeave RL Rollouts Hot-Loads Policy Weights Into Live Deployments, ~15x Faster Than Redeploys — _ScottCondron · 2026-10-07