AMD to acquire Taalas, whose HC1 chip etches Llama 3.1 8B into silicon at ~17,000 tokens/s

acmoytoy · x · 2026-09-15

AMD announced a definitive agreement to acquire Toronto-based Taalas, which takes a radically different approach to inference: instead of loading weights from HBM, the model's weights and dataflow are burned directly into the transistors — the chip is the model. Its first chip, HC1, runs Meta's Llama 3.1 8B at roughly 17,000 tokens per second per user on a single chip, punching through the memory wall that centralized AI labs have relied on. Commenters argue this points to powerful AI running on battery-powered devices in your pocket, undercutting narratives of centralized control over AI compute.

Original post →

More from Infra

Infra channel →