OpenAI Engineer: Anthropic's Tokenizer Change Snuck a ~30% Cost Hike into Opus 4.7
stevenheidel · x · 2026-09-04
OpenAI engineer Steven Heidel amplifies Maxime Labonne's finding that Anthropic snuck a 30% cost increase into Opus 4.7 by making its tokenizer less efficient — same text, more tokens, same listed price. He adds that 'tokens' aren't a standard unit across providers or even models from the same provider: 3.8 Flash looks 13x cheaper per token, but Astra can be cheaper per task because it's far more efficient. Measure costs per task, not per token.
More from Infra
- PyTorch AOTI Backend Speeds Up NVIDIA HSTU Inference by 1.14x–1.28x — PyTorch · 2026-09-04
- Real-Time Local 3B Model Demo on a ~$1800 Asus Zenbook, Qualcomm NPU Credited — draginol · 2026-09-04
- WCH's $3.30 Dual-Core RISC-V MCU Packs USB 3.2, 400MHz, and Ethernet — yacineMTB · 2026-09-04
- Inference startup insider: "we just resell NVIDIA GPUs" — VCs question the moat — firstadopter · 2026-09-04
- Leak claims GPT-6 Astra trained on 100,000+ GPUs at OpenAI's Stargate site — BLUECOW009 · 2026-09-04
- Users dispute credit burn; provider says KV cache was always on, scaling across providers — arthurcolle · 2026-09-04