AMD open-sources TokenSpeed-kernel: Triton as a semantic contract for multi-silicon LLM inference
zhyncs42 · x · 2026-10-02
AMD published an open-source subsystem, TokenSpeed-kernel, tackling backend complexity in multi-silicon LLM inference:
- Layered API design: the kernel-runtime interface stays platform-agnostic; the runtime owns model execution, scheduling metadata, page tables and routing state, while the kernel layer owns operator APIs, backend registration and high-performance implementations.
- Triton as a semantic contract: every kernel must ship a portable Triton baseline — increasingly valued not just for portability but as a numerical reference that precisely captures tensor semantics, enabling aggressive specialization (e.g., Gluon kernels with low-level scheduling and layout control). The philosophy: "languages as contracts, compilers as verifiers."
- Results: using GPT-OSS as a demo, both AMD and NVIDIA paths call the same public APIs; AMD's GPT-OSS 120B reaches top-tier performance via Gluon kernels, showing the layering doesn't sacrifice backend performance.
More from Infra
- Toshiba to invest ¥60B to double AI data center HDD capacity by fiscal 2027 — zephyr_z9 · 2026-10-02
- HeteroFold Enables Prefill-Free Cross-Family KV Cache Transfer, 10.7x Faster at 32K Context — UniversityofSouthernCalifornia · 2026-10-02
- Zuckerberg: Multi-Gigawatt Training Clusters Can Brute-Force Their Way to AGI — rohanpaul_ai · 2026-10-02
- Detect LLM hallucinations in 1.3µs on CPU — but 120B models hallucinate with unanimous false certainty — More_Slide5739 · 2026-10-02
- CodexBar: open-source menu bar app showing AI coding quotas (22k stars) — lxfater · 2026-10-02
- Micron earnings show HBM still dominates AI memory, bit growth solid through 2028 — AccBalanced · 2026-10-02