XGEN-JING: Xpeng's Open Egocentric World Model on MiniMax H3, Runs on Six H100s
AdinaYakup · x · 2026-09-21
XGEN-JING is Xpeng XGEN's egocentric interactive experience model built on MiniMax-H3, generating first-person video and audio for navigation, object interaction, and dialogue from actions, reference images, and observation history. Key details:
- Keyboard (WASD) camera control across real and imagined scenes
- Text-guided object interactions and character conversations with joint audio-video generation
- Up to 5 character/object/scene reference images; explore different actions from one starting point
- Current release: JING-Flash-v1 with 4-step bidirectional inference, inference code, examples, and prompt skills; causal model and technical report coming soon
- Validated on six H100s: one for text encoder, one for VAEs, four for sequence-parallel DiT; FlashAttention-4 default, Python 3.12 + SGLang runtime
Related event: XPeng Open-Sources First World Model XGEN-JING, Runs on 6 H100s(2 posts)→
More from Embodied
- Apple wins the AI infra lottery: M5 Ultra packs 512GB unified memory for local AI — Hesamation · 2026-09-21
- DeliveryGym: adaptive UE5 RL environment boosts Qwen3-VL-4B delivery earnings 54.3% — Lianhuiq · 2026-09-21
- Odyssey-3: one frozen world model drives humanoids, cars, drones, and games — thione · 2026-09-21
- Train a balance bot with PPO: Reinforcement Learning for Robotics Part 2 — ShawnHymel · 2026-09-21
- Google's $899 Googlebook bets you'll buy a new laptop just for Gemini — TechCrunch AI · 2026-09-21
- Archtyp AI CEO: Robots must be predictable in ways humans are not — xmercury_one · 2026-09-21