Tencent Releases Embodied Multimodal Foundation Model RxBrain
jacek2023 · reddit · 2026-07-15
This HF page introduces Tencent's Hy-Embodied-RxBrain-1.0, a unified multimodal foundation model designed for embodied cognition.
What It Does
- Embodied Understanding & Reasoning: Performs Q&A and chain-of-thought reasoning on images and multi-frame videos.
- World State Prediction: Imagines what visual frames will appear next in the physical world based on actions.
- Joint Sub-goal Planning: Breaks tasks into steps, simultaneously outputting the "next action" and a "target image" for each step.
Architecture Highlights
- Employs interleaved generation: alternately generating text reasoning and imagined frames within the same autoregressive sequence.
- Uses a single <Image> token to decide when to enter the "imagination" phase.
- The backbone is a unified Mixture-of-Transformers (MoT) with approximately 6.2B parameters.
- The image imagination component decodes into the latent space of a frozen FLUX VAE via a flow-matching image head.
Design Philosophy Emphasized
Rather than splitting understanding, generation, and planning into separate towers, it attempts to couple "what to say" and "what the world should look like" within the same sequence.
More from Embodied
- Tesla expands Robotaxi rides to seven areas, including new Orlando and Tampa zones — elonmusk · 2026-07-22
- Hands-on robotics workshop on Saturday may be the last in-person session before August — StewartalsopIII · 2026-07-22
- NVIDIA pitches World Foundation Models as a way to scale physical AI data generation — MonaJalal_ · 2026-07-22
- RoboMME Podcast Preview: Benchmarking Memory for Robotic Policies — chris_j_paxton · 2026-07-21
- Gritt says an 8-person crew now installs 3,000 to 4,000 solar panels a day — HaktanSuren · 2026-07-21
- A helium-powered flying robot whale aims to be a quiet companion pet — chris_j_paxton · 2026-07-21