User benchmarks claim OpenAI quietly cut model speeds to ~60% to stretch usage limits
bdsqlsz · x · 2026-10-01
A user benchmarked output throughput (total output tokens over 10 requests ÷ total elapsed time, including thinking tokens) and claims OpenAI has slowed its models to make usage quotas last longer.
- Astra, 5.6Sol, and Luna all measured at roughly 60% of their previous speeds.
- GPT-6.1 sol reportedly dropped to about 15 tok/s.
- The quoted tweet jokes that halving speed means "a day's quota now lasts a day and a half."
Unofficial user measurement, not confirmed by OpenAI.
Related event: Users Claim Codex Subscription Throttled to a Third of API Speed(2 posts)→
More from Models
- Hands-on GPT-6 Astra evals: big agentic gains, but ARC-AGI-3 scores swing wildly by harness — No-Soil-5789 · 2026-10-01
- Gemini 4 argon vs GPT 6.1 sol: same-prompt test shows starkly different outputs — iamfakhrealam · 2026-10-01
- Keep Claude on medium reasoning effort — high levels burn tokens and can hurt quality — intellectronica · 2026-10-01
- Technion's MIST stress test finds irrelevant images shift 20% of VLM judge labels regardless of content — Technion · 2026-10-01
- Reported Gemini 4 Argon touts 1M-token output window — leap or hype? — minimanishtic · 2026-10-01
- ChatGPT Users Report Older Chats and Projects Failing to Load, Fearing Data Loss — yaxir · 2026-10-01