Baseten Inference Engineering Masterclass: Quantization & Optimization
philipkiely · x · 2026-08-04
The Latent Space podcast invited key members from AI infrastructure company Baseten, which recently raised a $1.3 billion Series F, to dive deep into the technical core of Inference Engineering.
Key topics covered:
- Essence of Inference Engineering: Exploring how to turn trained model weights into fast, reliable, and scalable products.
- Performance Optimization: Detailed breakdowns of techniques like quantization and speculative decoding, explaining how quantization errors can cancel out to boost throughput.
- Industry Status: Noting that inference teams are still finding 20% to 200% performance gains.
- Video Generation Bottlenecks: Analyzing the quadratic attention wall faced by video generation compute.
- Model Self-Optimization: Sharing how the GLM-5.2 model helped rewrite and optimize the GPU kernels used to serve itself.
Related event: Baseten Team Shares Insights on AI Inference Optimization(3 posts)→
More from Infra
- Wan 2.1 now runs locally on supported Samsung phones via Saient Quartz — SaientAI · 2026-08-04
- NVIDIA pitches agentic commerce for retail, with merchant-controlled checkout and pricing — nvidia · 2026-08-04
- Stripe Projects lets AI agents add hosting, auth, databases, and billing from the CLI — jeff_weinstein · 2026-08-04
- Multi-agent workflows can burn billions of tokens unless you control duplication — HaktanSuren · 2026-08-04
- Next.js 16.3 cuts dev RAM by 90% and adds docs for coding agents — cramforce · 2026-08-04
- TokTier speeds up agent serving with exact stateful tokenization and stable-boundary repair — omarsar0 · 2026-08-04