Free Zoom meetup: disaggregated speculative decoding on d-Matrix chips plus inference engine tuning
cfregly · x · 2026-09-21
- Chris Fregly's AI Performance Engineering community is hosting a free Zoom meetup Monday 9am PT with two inference-performance talks.
- Talk 1: Tom St. John, Head of Applied Research at Gimlet Labs, demos disaggregated speculative decoding on d-Matrix chips via Gimlet Cloud.
- Talk 2: Makora's Head of Engineering demos tuning inference engine performance for the latest models and workloads.
- Resources shared: the ai-performance-engineering GitHub repo, an O'Reilly book on systems performance engineering, a YouTube channel and a free DeepLearning.ai GenAI course.
More from Infra
- How do you test LLM provider failure in production? A Reddit discussion — Rama_Surasani_ · 2026-09-21
- Run Flux 2 Dev (30B) locally with block offloading, mix models for T2I and editing — Altruistic_Heat_9531 · 2026-09-21
- Jensen Huang stands by $3-4T AI infrastructure market forecast — emmanuelvivier · 2026-09-21
- MiMo near-SOTA on DeepSWE with just ~$2.6M RL run: will data cost more than training? — my_cat_can_code · 2026-09-21
- Program-as-Weights: 0.6B model matches Qwen3-32B prompting with 1/50th memory, runs locally — yuntiandeng · 2026-09-21
- Personal AI Agents Are the Biggest Driver of the Sudden NAND Demand Surge — zephyr_z9 · 2026-09-21