Help: How Should Text Tags Be Written for Training WAN 2.2 Video LoRA?

Overall-Reporter-440 · reddit · 2026-08-03

A user encountered difficulties while using Musubi Tuner to train a LoRA for the WAN 2.2 image-to-video model. Following standard understanding, they detailed the actions in the video within the corresponding .txt files (e.g., "A man throws a tomato at an old man, splattering on his face, causing him to laugh") and included a trigger word.

However, after 1,000 training steps, testing in ComfyUI showed virtually no difference whether the LoRA was loaded or not. They are asking the community if there is a mistake or misconception in how these text tags are formulated.

Original post →

More from Multimodal

Multimodal channel →