AI Token usage up 1139%, costs down 55% amid caching shift
AccBalanced · x · 2026-08-18
Data shows weekly token usage surged 1139% year-to-date while cost per token dropped 55%. Key drivers include a shift in model mix towards cheaper frontier models and the increased use of cached tokens in agent workflows, which cost about 1/5th of uncached rates.
More from Infra
- Falcata: CUDA-native GBDT rebuilds training loop, 14x faster than LightGBM — srchvrs · 2026-08-18
- Brex Data: 14 of Top 25 Fastest-Growing Vendors Are AI Infrastructure — AccBalanced · 2026-08-18
- NVIDIA Boosts Local AI with Unsloth Integration and llama.cpp Optimizations — danielhanchen · 2026-08-18
- Text Watermark Detection Does Not Require Rerunning the LLM — rasbt · 2026-08-18
- Hot Chips 2026 preview: Focus on AI memory architectures and RISC-V evolution — AccBalanced · 2026-08-18
- Help: Configuring Ryzen Mini PC for Maximum LLM Inference Speed — Crafty-Sell7325 · 2026-08-18