POV Image Generation Emerges in Muse and Nano Banana, But Open-Source Models Fail

PangolinAdmirable881 · reddit · 2026-09-14

A Reddit user found that Meta's Muse Image and Nano Banana can generate images from a specific person's point of view given an input photo — a novel capability that's too costly to run at scale. Their open-source attempts all failed: Qwen Image Edit, FLUX.2 9B Base, and Hunyuan Image 3.0 Instruct couldn't produce POV shots; Gemma 4, asked to describe what the person sees, described the person instead; and Wan 2.2 TI2V couldn't shift the camera perspective. The thread solicits models or prompting tricks to crack POV generation with open tools.

Original post →

More from Multimodal

Multimodal channel →