Astribot launches Lumo-2 with real-robot demos

Astribot unveiled its second-generation embodied foundation model Lumo-2 ahead of WAIC, while also releasing the AgentPhilia agent and 20+ real-robot videos. Across the posts, the main point is that Lumo-2 is built around predicting task-relevant future physical states before acting, a design presented as more aligned with real robotic execution than generating full future videos.

Model design and training

According to the posts, Lumo-2 is a 4B robot foundation model based on a latent world-action model. It performs implicit predictive reasoning in latent space and models only changes useful for the task, such as object motion, contact changes, or when pouring has finished. To reduce ambiguity from a single frame, the model adds a short action-history buffer so it can infer which stage of a task it is in before deciding the next action. Its training is described as a three-stage process: first aligning visual changes with robot actions; then giving action codes semantic meaning through vision and language; and finally jointly training on VLM, video, and robot data to align actions, dynamics, vision, and language.

Performance and inference efficiency

Posts say Lumo-2 clearly outperforms Lumo-1 on multiple embodied reasoning tasks while remaining competitive with models aimed mainly at vision-language benchmarks. In one task-reasoning experiment, accuracy was 43% when using only the first frame, 94% with 5 observations, and 90% with only the first frame plus compact implicit dynamics; this is used to argue that the hidden state preserves key temporal information. For inference, Lumo-2 uses block-wise decoding, generating several weakly correlated action tokens at once instead of producing 32 action tokens one by one. According to the posts, it reaches 7.7 Hz on a single NVIDIA RTX 5090.

Real-robot demos and reactions

The released demos cover 22 complex household capabilities, including dual-robot collaboration, flipping an egg, weighing, cleaning a coffee grounds bowl, grinding beans, making coffee, mixing drinks, catching a rolling ball, tying a bow, zipping, and folding clothes. @CyberRobooo described the batch as among the stronger public real-robot demonstrations currently available, arguing that the value lies not only in execution quality but also in the display of physical understanding, temporal reasoning, and manipulation generalization.

2026-07-17 ~ 2026-07-18 · 10 related posts