Baseten's Inference Engineering Masterclass: Turning Model Weights into Production Apps
yenkel · x · 2026-08-08
The Latent Space podcast featured Philip Kiely and Ali Taha from Baseten for a deep dive into the core practices of Inference Engineering.
- Rise of Inference Engineering: Barely a distinct category three years ago, it is now one of the most critical disciplines in AI. It tackles the core problem: how to turn trained model weights into a fast, reliable, and affordable application at scale in production.
- Open Weights Debate: The episode contextualizes the current industry debate around open weights, with Ali recently publishing a viral, in-depth breakdown of the Kimi K3 modeling code.
- Free Ebook Release: Based on his AI Engineer talk and practical experience, Philip has authored a definitive guide to inference engineering, now available for free and popular across the SF tech scene.
More from Infra
- Running Qwen 3.6 27B on RTX 5090: 40 t/s at 262k Context in llama.cpp — Gargle-Loaf-Spunk · 2026-08-08
- Red Hat Releases New DSpark Models, Boosting vLLM Inference Speed by 4x — vllm_project · 2026-08-08
- Together AI Releases Interactive Diagrams Explaining LLM Inference and Quantization — zainhas · 2026-08-08
- Nscale Claims $51B Contracted Revenue Ahead of IPO, Faces Industry Skepticism — nathanbenaich · 2026-08-08
- vLLM and NVIDIA Achieve Over 25K TPS/GPU for Qwen3.5 on GB200 Systems — AccBalanced · 2026-08-08
- Prediction: Google Will Primarily Be a TPU Producing Business in a Decade — BorisMPower · 2026-08-08