Pinterest Spent 3 Months Hunting a Network Bug Caused by an Unused Container

arpit_bhayani · x · 2026-08-04

Pinterest's AI training jobs on Kubernetes kept crashing due to intermittent network loss. System logs initially pointed to AWS network driver resets—a self-healing mechanism triggered when a network thread is starved of CPU time for 5 seconds.

After usual troubleshooting failed, the team spent three months discovering that the culprit hogging CPU resources was a container they weren't even using.

Related event: Pinterest Finds Unused Container Caused GPU Crashes(2 posts)→

Original post →

More from Infra

Infra channel →