Pretraining compute share to drop from 67% to 7% as RL and inference take over, says SemiAnalysis
FinanceYF5 · x · 2026-09-17
SemiAnalysis estimates OpenAI and Anthropic's compute mix is flipping: pretraining's share of compute will fall from 67% to 7%, while post-training/RL jumps from 5% to 55% and inference reaches 38%. The next hardware battleground isn't raw compute but memory bandwidth.
Related event: Pretraining compute share to plummet from 67% to 7%: SemiAnalysis(2 posts)→
More from Infra
- Tencent open-sources FlexKV distributed KV cache for LLM inference, cutting TTFT by up to 70% — Roger_M_Taylor · 2026-09-17
- NVIDIA releases NVFP4 quantized DeepSeek-V4.1-Flash on Hugging Face — TheZachMueller · 2026-09-17
- Optimization mined via Bittensor competition lands in vLLM, boosting Qwen3 throughput ~4% — const_reborn · 2026-09-17
- Early vLLM PR adds Jev-like structured generation for DiffusionGemma, only 2x endpoint latency on a DGX Spark — generativist · 2026-09-17
- New inference engine Atlas debuts, redditor says it beats llama.cpp on Strix Halo — einthecorgi2 · 2026-09-17
- Keeping vLLM's prefix cache warm between agent turns: an engineering guide — bolts98 · 2026-09-17