API Cached Token Pricing Shifts O(n²) Attention Costs to Users

AccBalanced · x · 2026-07-07

A developer pointed out that API pricing based on "cached tokens" essentially allows OpenAI to pass the O(n²) computational cost of the attention mechanism onto the caller. This is an economic observation on the current billing models of mainstream LLM APIs.

Original post →

More from Models

Models channel →