Shopify shares 'Model Optimization Flywheel' for self-improving LLMs in production

MParakhin · x · 2026-09-01

Shopify presented their 'Model Optimization Flywheel' methodology at ICML 2026, designed to turn frontier LLM behaviors into faster, cheaper, and continuously improving production systems. The flywheel starts with LLM-as-judge evaluators grounded in human labels. Using Tangle workflows, they optimize system prompts, collect data from A/B traffic, and distill smaller models via SFT, on-policy distillation, and GRPO. These models replicate or exceed frontier behavior at lower cost. After deployment, the loop continues by healing low-scoring conversations with stronger models. This process has reduced costs and latency while improving quality.

Original post →

More from coding & agent

coding & agent channel →