XGEN-JING trends on Hugging Face: egocentric world model with joint audio-video generation
XGENlabs · hf · 2026-09-20
XGEN-JING from XGENlabs is trending on Hugging Face. It is an image-text-to-video model positioned as a world model with an egocentric (first-person) focus, and notably supports joint audio-video generation.
The model ships with diffusers and safetensors support, handles both English and Chinese, and is available for download on Hugging Face.
More from Multimodal
- Single HTML file adds weighted prompt randomization and consistency locks for ComfyUI — klartreumer · 2026-09-20
- GPT renders Opus's poetic imagining of Claude: a timeless oracle with fourteen eyes — repligate · 2026-09-20
- YuE2 tips: KSampler ~16x faster, CFG sweet spots, LoRAs and 900s song fix — Diligent-Rub-2113 · 2026-09-20
- AI Short Film 'It Happens to Every Wizard' Goes Viral on Reddit — Sweet_Composer_1323 · 2026-09-20
- Original AI Sci-Fi Short Film 'Slug' Asks What If Brainpower Could Be Sold — Luna_Nguyenn · 2026-09-20
- One GPT Image 2.5 prompt turns any photo's lower half into an office-supply collage illustration — Fresh-Resolution182 · 2026-09-20