User says image models ignore reference photos unless outfits are spelled out
Ok-Star-6755 · reddit · 2026-07-23
A Reddit user says image generation no longer seems to respect the reference photos they provide. They describe making small story scenes and finding that the model ignores outfit references unless they spell the clothing out explicitly.
The post is a concrete complaint about multimodal generation workflow: reference images appear to carry less weight than expected, so users may need to over-specify visual details in text to get consistent results.
More from Multimodal
- A terse French reply says the identity-law backlash is intentional — IgorCarron · 2026-07-23
- Facetnoir-style image generations shared with a reusable sref code — OVolosin82152 · 2026-07-23
- Midjourney images show fashion portraits and a stylized Ferrari render — Salmaaboukarr · 2026-07-23
- WAN 2.2 user asks whether video motion can be limited to one masked region — TekeshiX · 2026-07-23
- Seedance 2.0 gets Japanese dance videos right, with a funny ending — anthara_ai · 2026-07-23
- Storyboard-first workflow makes AI dance videos far more consistent — anthara_ai · 2026-07-23