Taalas bakes Llama 3.1 8B into custom silicon, hitting 15,000 tokens per second

npew · x · 2026-09-04

Taalas has burned Llama 3.1 8B directly into custom silicon, demonstrating generation at 15,000 tokens per second. Former OpenAI engineer Nathan Peter shared the demo and speculated about near-instant, Astra-class intelligence running on future phone chips. A notable datapoint for inference ASICs.

Original post →

More from Infra

Infra channel →