Taalas hits ~14,000 tokens/sec live, 140x ChatGPT, by baking models into silicon

iamfakhrealam · x · 2026-08-31

Emad Mostaque demonstrated Taalas live: 100 tokens/sec for ChatGPT vs 14,000 tokens/sec for Taalas — roughly 140x faster. The approach works by baking the model directly into the silicon itself, achieving extreme inference throughput through hardware-level model deployment.

Related event: Taalas demos 14,000 tokens/sec inference, ~100x faster than ChatGPT(3 posts)→

Original post →

More from Infra

Infra channel →