Tsinghua and Infinigence open-source APXInf, cutting embodied model latency 10.7x on Thor

机器之心 · wechat · 2026-09-29

Tsinghua University, Infinigence (无问芯穹) and Shanghai Jiao Tong University open-sourced APXInf, an on-device inference engine for embodied models. Without modifying π0.5, full-stack optimization cuts inference latency from 278ms to 26ms on Jetson Thor (FP8) — a 10.7x speedup reaching 38.46Hz real-time control — while LIBERO-10 success rate stays at 92.2%, matching the reference implementation. On Orin, optimization dropped latency from 1300ms to 119ms (10.9x). APXInf uses model-family-specific execution paths, CUDA Graph capture, pre-allocated memory and a Rust-based runtime, and turns model onboarding/tuning into an agent-executable workflow so humans only set standards and acceptance criteria. Already adapted for π0-fast, GR00T and QwenDrive, with AMD and domestic chip backends planned. The tech inherits from Mizar, preinstalled on 10M+ Lenovo AI PCs.

Original post →

More from Embodied

Embodied channel →