Inference > Training: LLMs are served year-round but trained once or twice

prajdabre · x · 2026-08-18

The author argues that while making training fast is important, optimizing inference speed is even more critical. Since LLMs are trained only once or twice a year but served continuously, inference performance should be the priority.

Original post →

More from Infra

Infra channel →