POV Image Generation Emerges in Muse and Nano Banana, But Open-Source Models Fail
PangolinAdmirable881 · reddit · 2026-09-14
A Reddit user found that Meta's Muse Image and Nano Banana can generate images from a specific person's point of view given an input photo — a novel capability that's too costly to run at scale. Their open-source attempts all failed: Qwen Image Edit, FLUX.2 9B Base, and Hunyuan Image 3.0 Instruct couldn't produce POV shots; Gemma 4, asked to describe what the person sees, described the person instead; and Wan 2.2 TI2V couldn't shift the camera perspective. The thread solicits models or prompting tricks to crack POV generation with open tools.
More from Multimodal
- MiniMax Design plugs into Blender via MCP to drive AI video from 3D references — Hailuo_AI · 2026-09-14
- Seedance 2.5 turns a school morning into a stealth game with impressive timing — SimplyAnnisa · 2026-09-14
- Codex + Lux3D + Blender workflow claims 3D models in ~20 seconds — SarahAnnabels · 2026-09-14
- Pippit Launches 3D Director Studio: Direct Your AI Video Instead of Prompting It — SarahAnnabels · 2026-09-14
- Robot-Themed AI Music Video Using Suno and MiniMax Turns Heads on Reddit — Educational-Mode-429 · 2026-09-14
- Japanese creator builds motion graphics with MiniMax H3 in HailuoAI workflow — Hailuo_AI · 2026-09-14