From 15% to 90% GPU Utilization: Fix the Data Pipeline, Not the Model

AI Engineer · youtube · 2026-10-10

In a talk at the AI Engineer World's Fair, Tarun Sunkaraneni of Amazon AGI shows that multimodal training is often CPU- and IO-bound, not GPU-bound. Training a Qwen3-VL-style model on S3 images, the baseline pipeline spent 85% of time waiting on data, with GPU utilization at just 15%.

He fixes bottlenecks one by one: concurrency with asyncio plus Ray actors, prefetching, and zero-copy tensor transfer via Ray's object store. At production scale a network-card bottleneck emerges, solved by spreading workers across nodes and zero-copy reads for another 50% throughput — cutting wait time from 85% to 20%. He also covers asyncio vs threads vs processes around the GIL, and why some fixes only pay off at scale.

Original post →

More from Infra

Infra channel →