Taalas hits ~14,000 tokens/sec live, 140x ChatGPT, by baking models into silicon
iamfakhrealam · x · 2026-08-31
Emad Mostaque demonstrated Taalas live: 100 tokens/sec for ChatGPT vs 14,000 tokens/sec for Taalas — roughly 140x faster. The approach works by baking the model directly into the silicon itself, achieving extreme inference throughput through hardware-level model deployment.
Related event: Taalas demos 14,000 tokens/sec inference, ~100x faster than ChatGPT(3 posts)→
More from Infra
- Increasing cache time saves 171M queries and 800GB egress daily — DanielLockyer · 2026-09-01
- Pinokio enables 1-click local AI model runs on PC — Art_If_Ficial · 2026-09-01
- How Much RAM is Needed for Qwen 3.8? A Reddit Discussion — Momsbestboy · 2026-09-01
- Over 4 Petabytes of Models and Datasets Uploaded to Hugging Face in One Week — victormustar · 2026-09-01
- Run Qwen 27B on 16GB VRAM: llama.cpp MTP mod adds 17% speed, more context — ea_man · 2026-09-01
- Analyst Weighs Long-Term Impact of Nvidia-MediaTek Deal on Broadcom — BenBajarin · 2026-09-01