XGEN-JING trends on Hugging Face: egocentric world model with joint audio-video generation

XGENlabs · hf · 2026-09-20

XGEN-JING from XGENlabs is trending on Hugging Face. It is an image-text-to-video model positioned as a world model with an egocentric (first-person) focus, and notably supports joint audio-video generation.

The model ships with diffusers and safetensors support, handles both English and Chinese, and is available for download on Hugging Face.

Original post →

More from Multimodal

Multimodal channel →