Infra Behind Krea 2: Tensor Cores and Crash Philosophy
AI Engineer · youtube · 2026-08-19
Gabriel Jorge Menezes shares the infrastructure behind training Krea 2 from scratch on thousands of GPUs. Key takeaways:
- Metrics: GPU utilization is a lie; track Tensor Core utilization. Build custom InfiniBand metrics.
- Hardware: Swap GPUs running >78°C to prevent throttling.
- Resilience: Crashes are inevitable; let it crash and retry, often succeeding for 24h.
- Storage/Scheduling: Aggressive checkpointing on a fast filesystem (1TB in <30s). Training outranks production, migrating inference gradually via fake K8s nodes.
More from Infra
- Inference Era Competition: Cloud Decisions Driven by Revenue per Watt — BenBajarin · 2026-08-19
- Code.Storage Opens Signups: Unlimited Git Infrastructure for AI Agents — dsp_ · 2026-08-19
- Lava lamps help secure internet data — dreamwieber · 2026-08-19
- Why cloud providers stopped auctioning compute: The $999 spot instance incident — lauriewired · 2026-08-19
- Engy launches permissionless TEE worker onboarding for confidential GPU compute — markjeffrey · 2026-08-19
- Apache Arrow Fix: Don't Trust Comments, Measure Actual Behavior — Abhishekcur · 2026-08-19