AMD user AI-codes ComfyUI nodes that speed up Qwen CLIP encoding 7.6x

kalexan94 · reddit · 2026-10-10

An AMD 7900XTX user used AI-assisted coding to write custom ComfyUI nodes that cut Qwen Image Edit 2.1 CLIP encoding from 51.3s to 6.7s (7.6x); full renders dropped from 66s to 21s, and prompt-only reruns from 62s to 18.5s.

How: the nodes patch the Qwen VL vision encoder in memory, swapping the vision patch-embed's Conv3d for a mathematically equivalent linear op (bit-identical output) with an A/B toggle. Separate Reference and Fast Text Encode nodes cache encoded reference images, so changing only the prompt skips re-encoding. The original node wastefully encoded every image twice (for positive and negative prompts); the mod encodes once.

Caveats: tested only on the author's machine (comfyui-rocm build), untested on NVIDIA; may also help other models using the Qwen VL vision tower (e.g. MiniMax H3).

Original post →

More from Infra

Infra channel →