One LLM already runs at 14,000 tokens/sec—frontier intelligence at 5,000 tok/s within 5 years?

dolo937 · reddit · 2026-09-02

The author speculates on inference speed as a paradigm shift: one LLM (chatjimmy.ai) already exceeds 14,000 tokens/sec, OpenAI's ultrafast mode serves 750 tok/s, and Minimax H3 generates video faster than humans can watch. They predict frontier intelligence at 5,000+ tokens/sec within five years, raising the question of a world where machine output outpaces human thinking and consumption.

Original post →

More from AGI Musings

AGI Musings channel →