User says image models ignore reference photos unless outfits are spelled out

Ok-Star-6755 · reddit · 2026-07-23

A Reddit user says image generation no longer seems to respect the reference photos they provide. They describe making small story scenes and finding that the model ignores outfit references unless they spell the clothing out explicitly.

The post is a concrete complaint about multimodal generation workflow: reference images appear to carry less weight than expected, so users may need to over-specify visual details in text to get consistent results.

Original post →

More from Multimodal

Multimodal channel →