Only 7% of training compute goes to pretraining? repligate probes what the figure means

repligate · x · 2026-10-10

repligate quotes a claim that only 7% of training compute goes to pretraining and asks whether it means 7% of a given checkpoint's training history, or 7% of a lab's total training compute — the former making sense since labs iterate on and abandon many post-training runs. The quoted tweet also worries RL pressure may override base models' learned goodness, though robust alignment could survive it.

Original post →

More from Models

Models channel →