Cerebras CTO's chip architecture deep dives—WSE-3, Hot Chips 34, Cornell lectures—barely get any views
blaizedsouza · x · 2026-09-09
A developer on a 365-day GPU programming journey rounds up Cerebras CTO Sean Lie's underrated public talks: a 2022 Cornell ML Hardware guest lecture covering architectural innovations, process technology gains, yield challenges, redundancy, lithography limits, cross-die wires, power/cooling and co-design; a WSE-3 core walkthrough; and a Hot Chips 34 deep dive into the memory system and fabric. Notably, these talks still have barely any views in 2026, showing how dispersed AI hardware knowledge remains.
More from Infra
- 27B model tops 100 tok/s on a MacBook via M5-tuned quantization and speculative decoding — teortaxesTex · 2026-09-09
- vLLM x AgentX: Full-Stack Optimizations for Real-World Agentic Serving — jfiance · 2026-09-09
- Desert Ant Labs introduces on-device intelligence for every product — Arcuru · 2026-09-09
- Speculative Decoding With Qwen3-30B-A3B Yields 1.5x Local Speedup, Up to 5x — Arindam_1729 · 2026-09-09
- Explainer: Speculative Decoding Speeds Up LLM Inference by ~100% — blaizedsouza · 2026-09-09
- Cosmos3 (64B) INT4 Quants Bring Local Image and Video Gen to Mac and CUDA — Formal-Swordfish-228 · 2026-09-09