InstinctFlash: open-source framework runs 5B robotics models in real time on Jetson Thor with 33.78x speedup

guanming0717 · hn · 2026-09-22

General Instinct released InstinctFlash, an AGPL-3.0 high-performance serving framework for robotics/VLA models. Runtime optimizations alone yield 1.2x–7.9x speedups on Jetson Thor; combined with a few-step distilled diffusion scheduler (25/50 steps reduced to 2/4), LingBot-VA reaches 33.78x. Across 50 Robotwin2.0 tasks (1,153 episodes per config), the 2/4-step setup hit 90.5% success vs 92.1% baseline. Optimizations include CUDA graph capture, cross-step KV/conditioning caching, mixed-precision GEMMs and specialized attention paths. It supports 8 VLA/world-action model families including pi0.5 and NVIDIA Cosmos Policy on RTX 4090/5090 and Jetson Thor, exposing accelerated checkpoints via a Python runtime or OpenPI-compatible WebSocket server.

Original post →

More from Embodied

Embodied channel →