ComfyUI Tutorial: Bypass Restrictions for Forced Image Captioning Using Modified QwenVL

Crack0saurus · reddit · 2026-07-31

A developer shared a method to bypass safety limits of the QwenVL image-to-text model in ComfyUI. While the original QwenVL has a strong text encoder, it rejects many images and hides its textual output.

By swapping the model directory with a modified version, Qwen3-VL-4b-Heretic, users can utilize an image-to-text node capable of describing literally anything. This hack requires no node code changes. The resulting captions can be used to extract poses, lighting, and style, acting as a substitute for a LoRA. The author notes that it takes about a minute to process, making it unsuitable for inline real-time workflows.

Original post →

More from coding & agent

coding & agent channel →