GPT Astra takes 38 minutes for 56k tokens: token efficiency is not compute efficiency
___Patrice___ · x · 2026-08-30
Observing that GPT Astra took 38 minutes to generate 56k tokens, the author notes that token efficiency does not equal compute efficiency. This suggests the model may be spending significantly more compute per token to produce fewer, higher-quality tokens. The trade-off raises questions about how this performance will be priced commercially.
More from Models
- Leak: Google Gemini 3.8 Flash coming soon with major quality boost — mark_k · 2026-08-30
- GLM-5.3-Flash Beats GPT-5 in Coding Arena at 26x Lower Cost — ccerrato147 · 2026-08-30
- Opus 5 criticized for jargon; author highlights LLMs' struggle to explain simple concepts clearly — antirez · 2026-08-30
- LongCat-Flash-Lite-Sparse and Qwen Uncensored Models Released in GGUF — LLMFan46 · 2026-08-30
- Qwen3.8-Flash-Next on 2x DGX Spark NVFP4: 50 t/s decode, 2,900 t/s prefill — -dysangel- · 2026-08-30
- Test: GLM 5.3 performs insane optimizations on Arctron AI — jasonkneen · 2026-08-30