SenseTime Turns Domestic Compute Into a Profitable Factory

新智元 · wechat · 2026-07-19

This in-depth article covers SenseTime's Agentic era AI infrastructure solution announced at WAIC, highlighting heterogeneous hybrid inference and the Token factory.

The article explains that LLM inference can be split into Prefill and Decode stages: the former relies on compute throughput, while the latter is memory bandwidth-bound. SenseTime's approach assigns different chip types to distinct roles: general domestic cards handle Prefill, and high-end cards manage Decode, streamlining the entire pipeline via system-level scheduling, compilers, caching, and network optimization. This solution has reportedly been deployed in real customer scenarios, adapting to models like GLM5.2, DeepSeekV4, KimiK2.7, MiniMaxM3, and SenseNovaV6.7.

Crucially, the commercial results indicate that hybrid inference clusters utilizing domestic chips have achieved profitability. This means revenue remains positive after covering depreciation, amortization, cabinet rentals, electricity, and personnel service fees. The average single-card MFU roughly doubled, peaking at a 152% improvement, and Token output per unit cost expanded 2.5 times. Furthermore, the daily average Token volume surged from 4000 billion earlier this year to an estimated 2.42 trillion, targeting 10 trillion by year-end, with expansions into energy storage, grid demand response, a national compute network, and overseas expansion in Hong Kong/Saudi Arabia and space-based computing.

Related event: SenseTime's Domestic AI Compute Turns Profitable(2 posts)→

Original post →

More from Infra

Infra channel →