OpenAI engineers: AI-found kernel optimizations cut GPT-5.6 Sol serving cost by 20%

TheTuringPost · x · 2026-09-12

OpenAI engineers Philippe Tillet and Matthew Ferrari explain how models now find bottlenecks and optimize inference themselves — kernel improvements alone cut GPT-5.6 Sol's end-to-end serving cost by 20%. The interview also covers which optimizations are wasted effort and how much models actually understand the systems they tune.

Original post →

More from AGI Musings

AGI Musings channel →