Minimax H3 Ref2VA Video Model Keeps Generating Unsettling Body Horror in Close-ups
fluvialcrunchy · reddit · 2026-09-22
A Reddit user reports the minimaxh3ref2va model (int8, no LoRAs, CK in the Ref2VA workflow) frequently partially undresses characters and produces disturbingly detailed body-horror anatomy in non-pornographic generations. Closer shots and words like "crotch," "legs spread," or "reveal" seem to raise the odds. No reliable workaround found yet; unclear whether official FL2VA/Ref2VF variants behave the same.
More from Multimodal
- Reddit Asks for Real-World LongCat-Video Inference Times on RTX 4090 to H100 — Clean_Extreme_3970 · 2026-09-22
- BUPT study: RoPE attention decay causes video diffusion models to violate physics — BUPT-CIST · 2026-09-22
- Tencent ARC's WorldCrafter adds implicit 3D-aware memory to video world models — TencentARC · 2026-09-22
- Grok 4.7 made this in Blender — demo shows the model driving 3D software — iamfakhrealam · 2026-09-22
- Kyutai releases Voice of Reason, a speech-native reasoning model hitting 77.1% on GSM8K — alexcovo_eth · 2026-09-22
- Qwen-Image local on a 24GB MacBook Pro takes 5-6 minutes per image — vista8 · 2026-09-22