Help: How Should Text Tags Be Written for Training WAN 2.2 Video LoRA?
Overall-Reporter-440 · reddit · 2026-08-03
A user encountered difficulties while using Musubi Tuner to train a LoRA for the WAN 2.2 image-to-video model. Following standard understanding, they detailed the actions in the video within the corresponding .txt files (e.g., "A man throws a tomato at an old man, splattering on his face, causing him to laugh") and included a trigger word.
However, after 1,000 training steps, testing in ComfyUI showed virtually no difference whether the LoRA was loaded or not. They are asking the community if there is a mistake or misconception in how these text tags are formulated.
More from Multimodal
- MiniMax H3 Tested: Generates High-Quality Videos in 2-4 Minutes — infroy28 · 2026-08-03
- Opinion: AI-Generated Disaster Movie Effects Will Surpass Live-Action Realism — JourneymanChina · 2026-08-03
- Same Prompt Image Generation Comparison: REVE 2.1 vs Meta AI — LudovicCreator · 2026-08-03
- NVIDIA's SANA-Video 2.0: Hybrid Attention Model Runs 720p on a Single RTX 5090 — mmowg · 2026-08-03
- Creating a Fantasy Movie Trailer Entirely with Hailuo AI — menhguin · 2026-08-03
- MiniMax H3 Open Weights Released: Community Tests Keyframe Interpolation — Hannibalj2ca · 2026-08-03