AMD user AI-codes ComfyUI nodes that speed up Qwen CLIP encoding 7.6x
kalexan94 · reddit · 2026-10-10
An AMD 7900XTX user used AI-assisted coding to write custom ComfyUI nodes that cut Qwen Image Edit 2.1 CLIP encoding from 51.3s to 6.7s (7.6x); full renders dropped from 66s to 21s, and prompt-only reruns from 62s to 18.5s.
How: the nodes patch the Qwen VL vision encoder in memory, swapping the vision patch-embed's Conv3d for a mathematically equivalent linear op (bit-identical output) with an A/B toggle. Separate Reference and Fast Text Encode nodes cache encoded reference images, so changing only the prompt skips re-encoding. The original node wastefully encoded every image twice (for positive and negative prompts); the mod encodes once.
Caveats: tested only on the author's machine (comfyui-rocm build), untested on NVIDIA; may also help other models using the Qwen VL vision tower (e.g. MiniMax H3).
More from Infra
- $500 ex-mining BC-250 cluster runs Qwen 35B at 145 tok/s with 256k context — Ok-Breadfruit-3523 · 2026-10-11
- SpaceX reportedly starting Terafab in December: a 100M sq ft chip megafactory — bennash · 2026-10-11
- OpenAI's Jalapeño chip trades kernel difficulty for memory bandwidth, betting on AI-written kernels — nrehiew_ · 2026-10-11
- Google Web AI lead Jason Mayes grew client-side JavaScript AI usage 2500x in 5 years — jason_mayes · 2026-10-11
- Dev in prod: running a newsletter side project on a cheap persistent VM — davidcrawshaw · 2026-10-11
- Zeeg: persistent VMs for agents are wrong, ephemeral sandboxes are the present — zeeg · 2026-10-11