GPT Astra takes 38 minutes for 56k tokens: token efficiency is not compute efficiency

___Patrice___ · x · 2026-08-30

Observing that GPT Astra took 38 minutes to generate 56k tokens, the author notes that token efficiency does not equal compute efficiency. This suggests the model may be spending significantly more compute per token to produce fewer, higher-quality tokens. The trade-off raises questions about how this performance will be priced commercially.

Original post →

More from Models

Models channel →