Dev predicts 10k-100k tok/s inference will make frontier models real-time robot policies

eigenron · x · 2026-09-07

eigenron predicts that within 1-2 years, fast inference on future chips like WSE-5/6 or next-gen Etched silicon will make general-purpose frontier models viable as real-time robot policies without task-specific training — roughly GPT-8 at 10k-100k tok/s per stream. He tested something recently that blew his mind and argues latency, not raw throughput, is the bigger bottleneck.

Related event: High-speed inference chips could turn frontier models into real-time robot policies(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →