Infini-AI launches Agentic Infra strategy to solve LLM inference bottlenecks
量子位 · wechat · 2026-07-20
At WAIC, Infini-Ai announced its "Front-store, Back-factory, One-center" Agentic Infra strategy to address compute supply-demand and inference cost bottlenecks. Its AgenticMaaS platform saw a 40x increase in daily Token calls compared to last year.
To tackle LLM inference pain points, Infini-Ai introduced the cross-cluster heterogeneous PD disaggregation architecture. By splitting Prefill and Decode phases across clusters and utilizing RadixCache and a 3-tier PDD architecture, it mitigates cross-cluster transmission latency. Tests show a 51.5% reduction in Time To First Token (TTFT) and a 37.5% drop in per-Token cost.
Additionally, the company released an intelligent cluster operations agent system, improving O&M efficiency by over 5x. Its compute distribution center has deployed 37,000P of compute, integrated 16 mainstream chips, and supported cross-domain reinforcement learning training with zero interruptions for a consecutive week.
Related event: Infinigence AI Unveils AgenticInfra Strategy at WAIC(2 posts)→
More from Infra
- PoLar: Dynamically Skipping or Looping LLM Layers for Efficient Inference — ttkciar · 2026-07-22
- SK Hynix CEO: Next Year Will Be the Worst Year in Industry's History from Supply Perspective — Beth_Kindig · 2026-07-22
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- Tech Giants Are Hiding $1.6T in AI Debt Using Enron's Trick — arto · 2026-07-22
- DeepSeek-V4-Flash tops out at 770 tok/s on one B300 in a vLLM batch test — Moreh · 2026-07-22
- NVIDIA starts shipping 102.4 Tbps Spectrum-6 switches for Vera Rubin AI factories — nvidia · 2026-07-22