MLLM-Guided Semantic Correction for Text-to-Video Generation
kwangmoo_yi · x · 2026-08-19
This paper introduces a training-free, interpretable mid-generation correction framework. By integrating multimodal large language model (MLLM) feedback directly into the diffusion sampling loop, it uses MLLMs to check denoising previews for semantic evaluation and injects corrective prompts. The method includes a Semantic Assessment Supervisor and a Semantic Modification Assistant, improving semantic alignment, visual fidelity, and temporal consistency without modifying model parameters.
More from Multimodal
- Reddit user uses AI to make a corny kids fantasy movie — MosskeepForest · 2026-08-19
- LTX-2.5 open weights demo generates impressive video results — tom_doerr · 2026-08-19
- Singularity Legend Vernor Vinge Created with ChatGPT Image 2.0 — DeryaTR_ · 2026-08-19
- LoRA released for video-image enhancing, upscaling, and restoring — CQDSN · 2026-08-19
- PNS paper: trainable neural subdivision augments Loop subdivision with bounded corrections — ssh4net · 2026-08-19
- Creating Vampire Hunter D Style Character with T2V and R2VA — SIR_NVAX_A_LOT · 2026-08-19