cHHillee: MFU tuning isn't a moat, but demand aggregation in ML infra is

cHHillee · x · 2026-09-22

cHHillee pushes back on the idea that kernel-level MFU gains are a significant advantage, arguing ML infra instead benefits from demand aggregation—the most efficient setup for a model like Kimi K3 might involve PD disaggregation plus wide EP. Patrick Toulme agrees GPU renting is a moat but disputes that MFU-tuning software is one, since future models may need far less tuning depending on how RSI plays out. cHHillee adds that large-scale RL posttraining infrastructure is resilient to future AI improvements, requiring substantial GPU-hours and token investment to replicate.

Related event: cHHillee and Toulme Debate Whether Tinker's Post-Training Service Has a Moat(8 posts)→

Original post →

More from Companies & People

Companies & People channel →