One LLM already runs at 14,000 tokens/sec—frontier intelligence at 5,000 tok/s within 5 years?
dolo937 · reddit · 2026-09-02
The author speculates on inference speed as a paradigm shift: one LLM (chatjimmy.ai) already exceeds 14,000 tokens/sec, OpenAI's ultrafast mode serves 750 tok/s, and Minimax H3 generates video faster than humans can watch. They predict frontier intelligence at 5,000+ tokens/sec within five years, raising the question of a world where machine output outpaces human thinking and consumption.
More from AGI Musings
- Researcher says AI-written papers are flooding literature, urges advisors to review drafts — dianarycai · 2026-09-02
- Cohere's research director: pre-training ROI is slowing, focus shifts to test-time compute — sarahookr · 2026-09-02
- Investors Have Poured $94.3B Into 119 'Neolabs' Betting on Research Before Revenue — WhatTheLJW · 2026-09-02
- Sarah Hooker: The Inference Era Is Breaking a Decade of GPU Design Assumptions — sarahookr · 2026-09-02
- Sarah Hooker: test-time compute and agentic workflows are forcing AI infra to change — sarahookr · 2026-09-02
- Criminologist on Hugging Face breach: multi-agent AI is the real safety frontier — sebkrier · 2026-09-02