Dev predicts 10k-100k tok/s inference will make frontier models real-time robot policies
eigenron · x · 2026-09-07
eigenron predicts that within 1-2 years, fast inference on future chips like WSE-5/6 or next-gen Etched silicon will make general-purpose frontier models viable as real-time robot policies without task-specific training — roughly GPT-8 at 10k-100k tok/s per stream. He tested something recently that blew his mind and argues latency, not raw throughput, is the bigger bottleneck.
More from AGI Musings
- Analyst speculates OpenAI's Astra is a World Model 3.0 after abandoning Sora — teortaxesTex · 2026-09-07
- Journal argues academia structurally selects against creativity — a decades-old consensus — davidmanheim · 2026-09-07
- Steering recursive self-improvement is 'close to nonsensical' amid alignment uncertainty — davidmanheim · 2026-09-07
- By Jensen Huang's own 2024 Stanford definition, AGI has arrived in some sense — Yuchenj_UW · 2026-09-07
- Should humans intervene in an alien civilization's path to its own singularity? — jachiam0 · 2026-09-07
- Nvidia CEO says 'AGI has arrived' and congratulates OpenAI — we_are_mammals · 2026-09-07