Alibaba's T-Head Semiconductor open-sources T-HeadSAIL, a CUDA-like stack for its Zhenwu AI chips serving 650+ customers

量子位 · wechat · 2026-09-25

Days after unveiling the Zhenwu V900 AI chip (claimed 3x the performance of M890) at Apsara Conference, Alibaba's T-Head Semiconductor expanded the open-source footprint of T-HeadSAIL, a CUDA-like software stack connecting PyTorch and other frameworks to its Zhenwu chips.

What's open now: PyTorch-for-sail framework adaptation, sailify source migration tool, Triton-for-sail kernel development, plus DeepGEMM-for-sail and FlashAttention-for-sail acceleration libraries, alongside SDKs, drivers, profiling and debugging tools. TensorFlow/JAX ports, an inference engine and communication components are still in progress.

Three pain points addressed: cutting chip-migration costs, letting developers inspect and tune acceleration libraries themselves, and pushing for Day0 support of new models — 39 quantized models (Qwen, DeepSeek, Kimi) already on ModelScope with 348k+ downloads.

Customer adoption: Zhenwu chips serve 650+ customers across 20+ industries. XPeng migrated autonomous-driving training to Zhenwu clusters; Ant Group completed inference adaptation for major frontier models on 810E/M890; Xiaohongshu built its own model-migration and kernel-optimization agent on the open code to speed generative recommendation deployment. Openness also resolves the tension of customers wanting custom kernels without handing over proprietary algorithms.

Why now: mature software plus real demand for secondary development; upstreaming adaptations to PyTorch, vLLM and Triton reduces the cost of maintaining a fork. It fits Alibaba's broader plan to invest RMB 380B+ over three years in cloud and AI infrastructure.

Original post →

More from Infra

Infra channel →