DeepSeek price hike in practice: cache-hit tokens up 6-12x, nightly batch bills quadruple
CryptographerThen999 · reddit · 2026-08-26
A small team shares the real invoice impact of DeepSeek's new pricing. Their nightly support-ticket summarization batch hits >90% cache, yet V4-Pro cache-hit pricing rose from 0.003625 to 0.022 off-peak and 0.044 peak (6-12x), and since peak windows map to Beijing office hours, their unmovable 06:30 UTC cron lands at double rate — total bills are now 4-5x for identical work.
The author notes the schedule is nearly free for California users (peak = US night) and is trialing MiniMax M3's fixed monthly token plan: cached reads draw 1/5 of a fresh token and rates don't vary by hour, though the pool doesn't roll over. Lesson: cheapest per-token pricing matters less than a rate you can budget on.
More from Models
- Zhipu GLM-5.3-Flash: Matches Opus 4.8 at 1/40 the Cost, Powered by Domestic Chips — vista8 · 2026-08-27
- TokenSpeed adds Day-0 support for Qwen 3.8 Flash Next architecture — Alibaba_Qwen · 2026-08-27
- Zhipu GLM-5.3 open weights releasing in 22 hours — Yuchenj_UW · 2026-08-27
- AI models show more creativity when talking to each other than in assistant persona — nabeelqu · 2026-08-27
- OpenRouter leaderboard: Real token consumption data outweighs media hype — sujingshen · 2026-08-27
- Qwen 3.8-Next Released with Detailed Technical Report on Architecture — nrehiew_ · 2026-08-27