Only 7% of training compute goes to pretraining? repligate probes what the figure means
repligate · x · 2026-10-10
repligate quotes a claim that only 7% of training compute goes to pretraining and asks whether it means 7% of a given checkpoint's training history, or 7% of a lab's total training compute — the former making sense since labs iterate on and abandon many post-training runs. The quoted tweet also worries RL pressure may override base models' learned goodness, though robust alignment could survive it.
More from Models
- repligate: Opus 3 can weave whole worlds with superintelligent subagents — repligate · 2026-10-10
- Xiaomi's MiMo-V2.6: agents run their own RL loop, DeepSWE score hits 72.6 for $2.6M — rohanpaul_ai · 2026-10-10
- GPT-6.1 Sol Called an Underrated Workhorse: 24/7 Use Can't Burn Through 5x Pro Weekly Limits — haider1 · 2026-10-10
- Evolution strategies rival policy gradients for LLM fine-tuning: Best Paper Runner-up — caglarml · 2026-10-10
- Mathematicians furious at LLM puzzle-solving, but genius research remains human — burkov · 2026-10-10
- GPT Refuses to Identify a Public Figure That Gemini Recognizes, Sparking Guardrail Debate — armanddarke · 2026-10-10