RAG and RL tool calling are bringing CPUs back into ML training, possibly needing dedicated CPU nodes
StasBekman · x · 2026-10-01
Stas Bekman highlights a shifting pattern in ML compute needs, updated in his ml-engineering repo's CPU chapter:
- CPUs traditionally only served DataLoader preprocessing while GPUs dominated.
- Around 2025, RAG workloads increased CPU usage via database queries.
- In 2026, AI tool calling in RL workloads is pushing CPU load further: executing programs to validate generated data/code, compiling, and running generated code sometimes exceeds co-located cores, requiring dedicated CPU nodes to avoid stalling GPUs.
- He also speculates CPUs may evolve to be more GPU-like and take on more offloaded work.
He asks the community for other unexpected CPU use cases.
Related event: RAG and RL tool calls put CPUs back in ML training(3 posts)→
More from Infra
- CoreWeave kicks off FC 2026 and unveils Forge, merging Weights & Biases, marimo and OpenPipe — altryne · 2026-10-01
- Comfy API goes live: deploy workflow JSONs as autoscaling GPU endpoints — PurzBeats · 2026-10-01
- Meta dodges billions in US taxes by classifying AI data centers as experiments — The Decoder · 2026-10-01
- E2B open-sources Embed, packaging full agent sandbox stack on a single node — badphilosopher · 2026-10-01
- ByteDance Seed finds phase sensitivity in chunked KV-cache compression, retrieval accuracy swings 40 points — ByteDance-Seed · 2026-10-01
- HPE/Broadcom Tomahawk trays for AMD Helios racks: six per rack, copper-heavy scale-up — BenBajarin · 2026-10-01