Self-Improving Pretraining: using post-trained models to pretrain better models
Ellen Xiaoqing Tan, Jack Lanchantin, Shehzaad Dhuliawala, Danwei Li, Thao Nguyen, Jing Xu, Ping Yu, Ilia Kulikov, Sainbayar Sukhbaatar, Jason Weston, Xian Li, Olga Golovneva
cs.CL, cs.AI, cs.LG
2026-01-29
A post-trained model rewrites pretraining suffixes and judges rollouts. Safety on a 1.4B model rises from 76.9 to 91.1; Llama-3-8B reasoning hits 3.2× direct RL after mid-training.
Standard training predicts the next token on raw web text, then tries to add safety, factuality, and reasoning in post-training. That is late. Patterns baked into the pretrained weights are hard to undo afterward. Deleting low-quality or unsafe documents does not fix the gap either: at inference the model still sees unsafe context, and pretraining never taught it to turn that context into a safe, higher-quality continuation.
Pretraining here is prefix-conditioned generation of a 128-token suffix. A fixed post-trained model rewrites the suffix and scores candidates. The pool is the original suffix, the rewrite, and rollouts from the policy being trained. Scores cover quality, safety, or factuality, prompted separately. Unsafe prefixes stay in the data; only the suffix is steered safe, so the policy still sees harmful context and has to learn the turn. The update is online DPO, with the top-scoring candidate chosen and the bottom one rejected. DPO is off-policy, so original suffixes and rewrites that the policy did not sample are valid training targets. Early on, the judge picks those references. Later it starts picking the policy's own rollouts.
For safety, the judge and the rewriter are GRPO-tuned from Llama3.1-8B-Instruct. A safe suffix scores only if it is copied. An unsafe suffix scores only after a safe rewrite. Factuality skips the rewriter and prompts GPT-OSS-120B, using the original suffix as the reference. The policy is Llama2-1.4B. Quality and factuality continual pretraining use SlimPajama, about 983k samples. Safety uses a harmful slice of RedPajama, about 257k. Continual runs take 2000 steps and up to 16 rollouts. From-scratch runs take 21k steps with a single rollout.
Thinking mid-training sits between pretraining and post-training, on Llama-3-8B. gpt-oss-120b inserts thoughts into DCLM and FineMath chunks where a reasoning step fits. Half of that data is SFT, with loss on both the original tokens and the thoughts. The other half is RL: the model writes a thought, then predicts the following text, and a judge scores only whether that prediction is close to the original suffix. The optimizer is DrGRPO. RL post-training then runs on DAPO-Math-14k, with answers checked by rules.
Generation quality is a pairwise judgment by GPT-OSS-120B on 128-token continuations. Llama Base against itself is 50%.
| Objective | Metric | Result |
| Quality | Win rate on standard prefixes / coherence | 86.3% / 87.9% |
| Factuality | Factuality average | 57.6 vs Llama Base 42.3 and continual next-token 44.0 |
| Safety | Safety average | 91.1 vs Llama Base 76.9 and RedPajama next-token 75.5 |
| All three | Standard-benchmark average | 50.8 / 50.5 / 49.1 vs Llama Base 47.6 |
Relative to Llama Base, factuality is up 36.2% and safety is up 18.5%. BoolQ, HellaSwag, MMLU and the other standard tasks move by about 1 to 3 points.
SFT on a single rollout collapses. Safety hits 99.5, the standard-prefix win rate falls to 2.0, and the standard-benchmark average falls to 29.5: the samples are meaningless strings that happen to look safe. SFT on rewrites reaches safety 86.5, with win rates of 52.7 on standard prefixes and 50.6 on unsafe ones. Online DPO of the original suffix against 16 rollouts is the run that gets a 73.6 win rate on standard prefixes, 77.7 on unsafe prefixes, and safety 91.1. Scaling rollouts from 1 to 16 improves quality, factuality, and safety. At 8 rollouts, a finetuned 8B judge scores 72.1 on quality win rate. Prompted GPT-OSS-120B scores 84.3.
From scratch, with one rollout, RF-NLL reaches a 32.4 win rate against Llama Base, versus 1.3 for from-scratch next-token prediction. Safety is 97.5 versus 85.2. 32.4 is still under the 50% tie, so this run has not caught a finished Llama Base.
Thinking mid-training is scored after RL post-training as pass@1 averaged over 16 samples, on GSM8K, MATH-500, Olympiad, AMC23, and GPQA-Diamond.
| Pipeline | Mid-train tokens | Average |
| Base, then RL | 0 | 0.1197 |
| SFT on raw text, then RL | 10.5B | 0.1645 |
| SFT on thoughts, then RL | 10.5B | 0.3480 |
| Thought SFT plus RL mid-training, then RL | 11B | 0.3837 |
0.3837 is 3.2× direct RL on the base model, and above the 0.1645 from raw-text mid-training. On the full pipeline, GSM8K is 0.7934 against 0.2172 for direct RL, MATH-500 is 0.4612 against 0.1038, Olympiad is 0.1593 against 0.0211, and GPQA-Diamond only moves from 0.2314 to 0.3220. Before any post-training, thought SFT plus RL mid-training already averages 0.3390, against 0.0475 for the base and 0.0999 for raw SFT. Extra SFT tokens do less: 7.8B to 10.5B of thought SFT moves the post-RL average from 0.3346 to 0.3480, while 8.7B that includes RL mid-training reaches 0.3785. On Qwen3-8B-Base, raw SFT drops the average from 0.3572 to 0.2802, and RL mid-training brings it back to 0.3660.
Rewriting and LLM-as-judge, usually post-training tools, now shape pretraining. Keeping the unsafe prefix and repairing the suffix teaches a steer, which deleting unsafe documents never teaches. Thinking mid-training makes think-then-continue a habit on ordinary web text, so post-training does not have to install that format from scratch.
When a strong post-trained model is already in hand, it can supervise a smaller student's earlier stages. The paper reports that the safety- and quality-tuned 1.4B beats Llama-3.1-8B Base, their check on a pure distillation account. Every step needs rollouts plus a judge call, which is slower than next-token prediction. The runs stop around a million documents and about 11B mid-training tokens.
The loop is not closed. Teachers are Llama3.1-8B-Instruct and GPT-OSS-120B. Students are Llama2-1.4B and Llama-3-8B. Nothing here trains a next generation from the student.
At eval time, generation quality, safety, and factuality are all scored by GPT-OSS-120B. That same model is the training judge for quality and factuality. The safety training judge is the finetuned 8B. Standard knowledge benchmarks barely move, so the gains sit next to the judge's preferences. The factuality judge treats the original suffix as ground truth, and web text is often wrong. Safety training does not buy factuality, or the reverse. A single run with all three rewards is not in the paper. Once safety is the default, tasks that need unsafe text get worse. The control-token switch mentioned in the paper is not in the main experiments.
The thought reward checks whether the predicted suffix matches the original, not whether the thought is valid. Thought length rises with the reward, and length hacking is not isolated. The 3.2× is mostly math. GPQA moves less, and post-training data is DAPO-Math only. Llama-3-8B base GSM8K is 0.0123, so the multiplier starts from a very low floor. On Qwen3 the gain nearly disappears. The two halves of the paper are not one pipeline.