Baseten masterclass says inference teams still find 20% to 200% speed gains
afurgs · x · 2026-08-04
This post points to a long podcast/masterclass on inference engineering with Baseten.
Key themes include:
- turning trained weights into a fast, reliable, affordable product is its own optimization problem
- inference teams are still finding 20–200% performance gains
- quantization, speculative decoding, and GPU-kernel optimization are central
- video generation hits a quadratic attention wall
- GLM-5.2 reportedly helped rewrite and optimize the GPU kernels serving GLM-5.2 itself
The thread also frames inference engineering as a now-critical discipline, with the broader AI infra stack benefiting from the shift toward serving and optimization.
Related event: Baseten Team Shares Insights on AI Inference Optimization(3 posts)→
More from Infra
- Wan 2.1 now runs locally on supported Samsung phones via Saient Quartz — SaientAI · 2026-08-04
- NVIDIA pitches agentic commerce for retail, with merchant-controlled checkout and pricing — nvidia · 2026-08-04
- Stripe Projects lets AI agents add hosting, auth, databases, and billing from the CLI — jeff_weinstein · 2026-08-04
- Multi-agent workflows can burn billions of tokens unless you control duplication — HaktanSuren · 2026-08-04
- Next.js 16.3 cuts dev RAM by 90% and adds docs for coding agents — cramforce · 2026-08-04
- TokTier speeds up agent serving with exact stateful tokenization and stable-boundary repair — omarsar0 · 2026-08-04