Reka unveils Rho-1, a 19B unified any-to-any model trained on 320 H100s
Reka AI has released a research preview of Rho-1, a 19B-parameter omnimodal model. A single neural network can understand and generate text, images, video, and even robot actions; it was trained from scratch using only about 320 H100 GPUs in roughly 3 months, with the team saying the compute investment is a tiny fraction of today's frontier models. This is seen as an attempt to build a unified omnimodal world model at low cost.
Confirmed
- Model specs and training cost: 19B parameters, trained from scratch on about 320 H100 GPUs over roughly 3 months; officially stated compute is far below the frontier-model average.
- Unified architecture claim: Unlike the current mainstream agentic pipeline approach (multi-agent pipelines + specialized model division of labor), Rho-1 unifies understanding and generation of text, images, video, and robot actions within a single network; the team says this is their first attempt to demonstrate that "unified omnimodal world models are the future of AI."
- Physical AI closed loop: Officials say predicting the camera's next frame and generating actions use the same set of weights, closing the loop from perception to action.
- Steerable world-model capability: Rho-1 can generate continuous video in real time and update trajectories mid-stream when new instructions arrive without cutting the feed. The team stresses that "real-time + steerable" is exactly what separates rendering a video from running a simulation—i.e., the model doesn't generate clips but an interactively drivable simulation.
Why it matters
- If the unified world-model route proves viable, it avoids the coordination costs of stitching multiple models together; Rho-1 was trained with far less compute than frontier models, offering a low-cost reference path for small and mid-sized teams.
- If real-time interactive simulation matures, it could serve robotics control, embodied AI, and other physical AI scenarios—the "rendering vs. simulation" framing highlights its fundamental difference from ordinary video generators.
This is currently a research preview; actual capability and stability await independent community verification.
2026-10-05 ~ 2026-10-06 · 8 related posts
Primary sources
- Reka releases Rho-1 research preview: a 19B omni model unifying text, images, video and robot actions — RekaAILabs ·
- Reka details Rho-1 training: from scratch on 320 H100s in ~3 months, a fraction of frontier compute — RekaAILabs ·
- Rho-1 streams continuous video in real time and re-steers mid-stream on new instructions, Reka says — RekaAILabs ·
- [source] Reka releases Rho-1 research preview: a 19B omni model unifying text, images, video and robot actions — RekaAILabs · 2026-10-05
- Rho-1 doubles as a steerable world model: real-time video you can redirect mid-stream — RekaAILabs · 2026-10-05
- [source] Rho-1 streams continuous video in real time and re-steers mid-stream on new instructions, Reka says — RekaAILabs · 2026-10-05
- [source] Reka details Rho-1 training: from scratch on 320 H100s in ~3 months, a fraction of frontier compute — RekaAILabs · 2026-10-05
- Reka releases Rho-1, a 19B omni-model trained from scratch on 320 H100s in 3 months — RekaAILabs · 2026-10-05
3 near-duplicate retellings: RekaAILabs · RekaAILabs · altryne