Switching models frequently invalidates prompt cache, causing double billing

daniel_mac8 · x · 2026-08-25

@Teknium advises against constantly switching models within a single session. Doing so invalidates the prompt cache on the new model, forcing you to repay the full input token price for the entire context. This is a fundamental aspect of inference, not specific to any one model. Avoid this unless the models are free.

Related event: Frequent model switching wipes prompt cache and doubles inference cost, Teknium warns(2 posts)→

Original post →

More from Infra

Infra channel →