Grok 4.5 Uses Significantly Fewer Tokens
XFreeze · x · 2026-07-09
According to a post, Grok 4.5 consumes far fewer output tokens on SWE Bench Pro compared to Claude Opus 4.8 max. It averages 15,954 output tokens versus 67,020 for the latter, roughly a 4.2x reduction. The author also reported a response throughput of 80 TPS, emphasizing lower inference costs, faster responses, and better scalability.
Related event: xAI Launches Grok 4.5: Coding and Agent Focus to Rival Opus(55 posts)→
More from Infra
- Dual R9700 on Asus X570 vs jumping to AM5: a ~€1000 local LLM upgrade dilemma — Mountain-Badger-5815 · 2026-09-07
- Tracking one tennis ball with GPT-6 burns 7.87M tokens — Scobleizer · 2026-09-07
- Huawei's Kirin 9050Pro uses logic folding to cut NPU power 66%, run 30B MoE on-device — APPSO · 2026-09-07
- Energy and the power grid, not chips, are the real bottleneck for AI at 400k-GPU scale — kimmonismus · 2026-09-07
- Signing TLS handshakes inside a TPM to protect machine identity — jedisct1 · 2026-09-07
- GLM 5.3 and Qwen 3.8 now run really well locally on single desktops — jasonkneen · 2026-09-07