AMD to acquire Taalas, whose HC1 chip etches Llama 3.1 8B into silicon at ~17,000 tokens/s
acmoytoy · x · 2026-09-15
AMD announced a definitive agreement to acquire Toronto-based Taalas, which takes a radically different approach to inference: instead of loading weights from HBM, the model's weights and dataflow are burned directly into the transistors — the chip is the model. Its first chip, HC1, runs Meta's Llama 3.1 8B at roughly 17,000 tokens per second per user on a single chip, punching through the memory wall that centralized AI labs have relied on. Commenters argue this points to powerful AI running on battery-powered devices in your pocket, undercutting narratives of centralized control over AI compute.
More from Infra
- NVIDIA: full-stack NIM tuning delivers 2.5x more concurrent users on Nemotron 3 Ultra — NVIDIAAI · 2026-09-15
- How Much Does Local LLM Inference Really Cost? A Dev Added an Electricity Calculator — giveen · 2026-09-15
- Hugging Face Rounds Up Which Open LLMs Are Best for On-Device Inference — NielsRogge · 2026-09-15
- Single Pure-C99 Inference Engine Runs Both BitNet Ternary and GGUF, No Python or CUDA — shifu_legend · 2026-09-15
- Dev weighs ChatGPT subscription via OAuth vs API pricing for a production RAG app — builtforoutput · 2026-09-15
- Cognichip launches ACI Enterprise: one engineer finishes chip front-end design in 10 days — kimmonismus · 2026-09-15