Taalas bakes Llama 3.1 8B into custom silicon, hitting 15,000 tokens per second
npew · x · 2026-09-04
Taalas has burned Llama 3.1 8B directly into custom silicon, demonstrating generation at 15,000 tokens per second. Former OpenAI engineer Nathan Peter shared the demo and speculated about near-instant, Astra-class intelligence running on future phone chips. A notable datapoint for inference ASICs.
More from Infra
- Nvidia A100 pre-training pipeline goes live on a 24/7 public livestream — wavefnx · 2026-09-04
- Spotify's Portal Cut Claude Code Token Usage by 90% With a Two-Mode Router — rseroter · 2026-09-04
- Vyact: Open-Source Desktop Workspace Unifying Local LLMs, RAG, and Browser Context — vyact · 2026-09-04
- At what context depth does KV quantization start to hurt? An F16 vs Q8/Q4 parity PoC — Slight_Analysis_5414 · 2026-09-04
- KV caching: the fundamental optimization behind autoregressive LLM inference — alec_helbling · 2026-09-04
- WSJ: Data centers are a "lottery win" for workers in Richland Parish, Louisiana — robleclerc · 2026-09-04