线上分享:用 d-Matrix 芯片做分解式投机解码与推理引擎调优
cfregly · x · 2026-09-21
- Chris Fregly 的 AI Performance Engineering 社区将于周一上午 9 点(PT)在 Zoom 举办两场推理性能分享。
- Talk 1:Gimlet Labs 应用研究负责人 Tom St. John 演示在 Gimlet Cloud 上用 d-Matrix 加速器实现分解式投机解码(disaggregated speculative decoding)。
- Talk 2:Makora 工程负责人演示如何针对最新模型与工作负载调优推理引擎性能。
- 相关资源:GitHub 仓库 ai-performance-engineering、O'Reilly 新书《Systems Performance Engineering》、YouTube 频道及 DeepLearning.ai 免费 GenAI 课程。
「Infra」频道最新
- 开发者呼吁:避免 cache miss 或是提升 Claude Code/Codex 用量上限最关键一招 — chaseleantj · 2026-09-21
- Reddit 实操讨论:生产环境如何测试 LLM 提供商故障 — Rama_Surasani_ · 2026-09-21
- 本地跑 Flux 2 Dev 全攻略:30B 模型配显存方案与模型组合玩法 — Altruistic_Heat_9531 · 2026-09-21
- 黄仁勋重申 AI 基础设施市场将达 3-4 万亿美元;Meta 遭生物识别诉讼 — emmanuelvivier · 2026-09-21
- 260 万美元 RL 就近 SOTA:数据成本会否超过训练成本? — my_cat_can_code · 2026-09-21
- Program-as-Weights:0.6B 模型本地性能媲美 32B 提示,内存仅 1/50 — yuntiandeng · 2026-09-21