SenseTime Turns Domestic Compute Into a Profitable Factory
新智元 · wechat · 2026-07-19
This in-depth article covers SenseTime's Agentic era AI infrastructure solution announced at WAIC, highlighting heterogeneous hybrid inference and the Token factory.
The article explains that LLM inference can be split into Prefill and Decode stages: the former relies on compute throughput, while the latter is memory bandwidth-bound. SenseTime's approach assigns different chip types to distinct roles: general domestic cards handle Prefill, and high-end cards manage Decode, streamlining the entire pipeline via system-level scheduling, compilers, caching, and network optimization. This solution has reportedly been deployed in real customer scenarios, adapting to models like GLM5.2, DeepSeekV4, KimiK2.7, MiniMaxM3, and SenseNovaV6.7.
Crucially, the commercial results indicate that hybrid inference clusters utilizing domestic chips have achieved profitability. This means revenue remains positive after covering depreciation, amortization, cabinet rentals, electricity, and personnel service fees. The average single-card MFU roughly doubled, peaking at a 152% improvement, and Token output per unit cost expanded 2.5 times. Furthermore, the daily average Token volume surged from 4000 billion earlier this year to an estimated 2.42 trillion, targeting 10 trillion by year-end, with expansions into energy storage, grid demand response, a national compute network, and overseas expansion in Hong Kong/Saudi Arabia and space-based computing.
Related event: SenseTime's Domestic AI Compute Turns Profitable(2 posts)→
More from Infra
- NVIDIA says Blackwell Ultra hit 1,648 TFLOPs per GPU on DeepSeek-V3 671B training — NVIDIAAI · 2026-07-21
- NVIDIA repeats its Blackwell Ultra throughput claim on DeepSeek-V3 671B — NVIDIAAI · 2026-07-21
- A shopping app demo ties OpenTelemetry, Dynatrace and Port into agentic ops — Pavan_Belagatti · 2026-07-21
- Nvidia Rubin is coming, pointing to the next AI compute platform — ezyang · 2026-07-21
- Tesla’s FSD v14 Lite is reportedly headed to 4 million older HW3 cars — MatthewBerman · 2026-07-21
- TSMC’s 3nm utilization reportedly tops 120% as AI demand drives a $190B capex cycle — tengyanAI · 2026-07-21