Paper: A Year in LLM Serving Analysis on 6.1B Requests Reveals Caching Insights

JiaZhihao · x · 2026-08-22

Based on a one-year production trace from Chutes.ai (6.1B requests, 315K users, 9,174 models), this paper analyzes the evolution of LLM serving workloads.

Key Findings:

Original post →

More from Infra

Infra channel →