Grok Imagine fails ~80% of the time on nylon mask consistency in I2V
dropdead90s · reddit · 2026-09-04
A Reddit user tested how video generation models handle nylon/pantyhose materials as masks: with Grok Imagine's image-to-video, the model tears a hole at the mouth or ignores the mask-over-mouth prompt roughly 80% of the time. The post asks whether others have seen similar behavior with SD-based models, highlighting a general weakness in occlusion and material consistency for current video generation.
More from Multimodal
- 'Sometimes I wish I live my life as AI': MiniMax H3 pixel-art girl goes viral — aziz4ai · 2026-09-04
- ICML 2026 oral talk live: motion attribution for video generation — cindy_x_wu · 2026-09-04
- MiniMax H3 Max Director lands on fal: realtime streaming video at $0.02/sec promo price — noahsolomon · 2026-09-04
- MiniMax H3 generates lifelike pixel-game-style video, upscaled with Magnific — MarioKrenn6240 · 2026-09-04
- Dev rebuilds Vine with H3 Max Turbo on fal: an infinite video feed that generates faster than you scroll — chrisfirst · 2026-09-04
- ICML oral Motive: first motion attribution framework for video generation, 74.1% win rate — cindy_x_wu · 2026-09-04