WorldFM Open Source: Real-Time Multi-View Diffusion from Target Poses
tom_doerr · x · 2026-08-14
A new project named WorldFM has been open-sourced on GitHub. It is a real-time multi-view diffusion model that generates images at new viewpoints given a single reference image and target camera poses.
The repository includes complete setup scripts and integrates submodules like HunyuanWorld-1.0, MoGe, Real-ESRGAN, and ZIM, allowing developers to quickly build the environment and run the inference pipeline.
More from Multimodal
- Opus 5 + Higgsfield One-Shots a 3D Platformer Game Entirely via AI Video Generation — TAbrodi · 2026-08-14
- LTX 2.5 Tested: IC LoRAs from Version 2.3 Remain Compatible — No-Property3068 · 2026-08-14
- Demo: Generating Realistic First-Person Guitar Playing with Luma and ElevenLabs — mrjonfinger · 2026-08-14
- Open Source Reel Video: Combines Subscriptions to Cut AI Video Costs to $50/Month — StepUpPrep · 2026-08-14
- Testing ComfyUI-H3-FaceRefine: A Node for Enhancing Video Face Details — Devajyoti1231 · 2026-08-14
- Flora Launches Fashion Studio to Streamline Creative Workflows — round · 2026-08-14