Cold start drops from 45min on vLLM to 3min with zml_ai
steren · x · 2026-10-12
Developer steren reports model server cold start going from 45 minutes on vLLM to just 3 minutes after switching to @zmlai. ZML founder @steeve explains the gain comes from recent compilation-time optimizations worked on around the zml/llmd release.
More from Infra
- OpenRouter token traffic explodes from 2T to 379T/month, open-weight models at 75% — Beth_Kindig · 2026-10-12
- ASUS RTX 5090 listed at €11,079 in Germany, only 3 units left in stock — janusch_patas · 2026-10-12
- Modular Claims Mojo Kernels Beat FlashAttention 4, Launches MAX Inference Framework — AI Engineer · 2026-10-12
- DeepSeek-V4.1-Flash on 2x DGX Spark: TP2 Patches Cut First Token From 33.8s to 4.9s — rez0__ · 2026-10-12
- Ed Zitron: AI needs $425B annual revenue to justify hyperscaler capex, $260B short — SumitGup · 2026-10-12
- DiffusionBear: MLX-powered macOS app runs FLUX.2 & Krea 2 on 16GB Macs — Puzzleheaded_Note739 · 2026-10-12