Grok Imagine fails ~80% of the time on nylon mask consistency in I2V

dropdead90s · reddit · 2026-09-04

A Reddit user tested how video generation models handle nylon/pantyhose materials as masks: with Grok Imagine's image-to-video, the model tears a hole at the mouth or ignores the mask-over-mouth prompt roughly 80% of the time. The post asks whether others have seen similar behavior with SD-based models, highlighting a general weakness in occlusion and material consistency for current video generation.

Original post →

More from Multimodal

Multimodal channel →