From 15% to 90% GPU Utilization: Fix the Data Pipeline, Not the Model
AI Engineer · youtube · 2026-10-10
In a talk at the AI Engineer World's Fair, Tarun Sunkaraneni of Amazon AGI shows that multimodal training is often CPU- and IO-bound, not GPU-bound. Training a Qwen3-VL-style model on S3 images, the baseline pipeline spent 85% of time waiting on data, with GPU utilization at just 15%.
He fixes bottlenecks one by one: concurrency with asyncio plus Ray actors, prefetching, and zero-copy tensor transfer via Ray's object store. At production scale a network-card bottleneck emerges, solved by spreading workers across nodes and zero-copy reads for another 50% throughput — cutting wait time from 85% to 20%. He also covers asyncio vs threads vs processes around the GIL, and why some fixes only pay off at scale.
More from Infra
- $500 ex-mining BC-250 cluster runs Qwen 35B at 145 tok/s with 256k context — Ok-Breadfruit-3523 · 2026-10-11
- SpaceX reportedly starting Terafab in December: a 100M sq ft chip megafactory — bennash · 2026-10-11
- OpenAI's Jalapeño chip trades kernel difficulty for memory bandwidth, betting on AI-written kernels — nrehiew_ · 2026-10-11
- Google Web AI lead Jason Mayes grew client-side JavaScript AI usage 2500x in 5 years — jason_mayes · 2026-10-11
- Dev in prod: running a newsletter side project on a cheap persistent VM — davidcrawshaw · 2026-10-11
- Zeeg: persistent VMs for agents are wrong, ephemeral sandboxes are the present — zeeg · 2026-10-11