Theory: model 'nerfing' may come from mixed heterogeneous inference hardware, not intent

michellechen · x · 2026-09-24

The author floats an unconfirmed theory: complaints that a model has been "nerfed" may actually stem from providers running inference on heterogeneous hardware (GPUs, TPUs, AMD, Inferentia), which produces varying output quality. She notes this lines up with nerfing appearing shortly after launch or during peak-time capacity crunch. Purely speculative.

Original post →

More from Infra

Infra channel →