Pretraining Still Matters: RL Is Just More Compute-Intensive, Not More Important
akbirthko · x · 2026-10-04
A brief but substantive exchange on the shifting balance between pretraining and RL. The claim: RL being more compute-intensive doesn't make pretraining less important — pretraining provides dense signal that RL can't replicate. The compute weighting difference reflects cost asymmetry, not a fundamental devaluation of pretraining.
More from Infra
- Samsung: HBM to consume 30% of DRAM wafer capacity by 2027, up from ~20% — Beth_Kindig · 2026-10-04
- Nebius up 160% vs CoreWeave's 10%: the deciding factor isn't revenue or backlog — Beth_Kindig · 2026-10-04
- US Data Center Construction Spending Surged 73% YoY to a Record $85B Annualized — FlorianGallwitz · 2026-10-04
- Developer's talk on running on-device AI praised for great insights — carrycooldude · 2026-10-04
- Cloudflare: pending I/O now keeps Durable Objects alive for agents without a connected client — irvinebroque · 2026-10-04
- Dev ships C-based local inference engine targeting tool calling on low-VRAM hardware — ZenZombie117 · 2026-10-04