Hanshu Tech unveils uHBM and uLPU inference architecture

新智元 · wechat · 2026-08-31

Hanshu Tech unveiled the uHBM® and uLPU™ inference architecture to address the weight transfer bottleneck in LLM Decode stages. By integrating Persistent MRAM with matrix-vector compute on a single die (Weight-Resident Compute), it achieves 24TB/s in-die read bandwidth, avoiding repeated weight transfers across storage interfaces.

Original post →

More from Embodied

Embodied channel →