Cross-architecture weight grafting merges image and video models with different backbones
SelectionNormal5275 · reddit · 2026-07-21
Cross-architecture weight grafting can merge image and video models with different backbones
This Reddit post describes an experimental method called Cross-Architecture Weight Grafting for transplanting small parts of one model into another even when their architectures and layer shapes do not match.
What the author tried
- Image models: Qwen-Image-2512 as the base, with Krea 2, Flux Dev 2, and Klein 9B as donors.
- Video models: LTX 2.3 Dev and LTX 2.3 Distilled 1.1 as bases, with Wan 2.2 low-noise as the donor.
Main findings
- The best result reported was Krea 2 → Qwen-Image-2512.
- Qwen Image → Krea 2 lost realism and looked weaker than the original Krea style.
- Flux Dev 2 → Flux Dev 1 loaded successfully but produced blurry images, suggesting the method needs more tuning.
- Flux 2 Dev → Klein 9B also loaded and ran.
- Early tests merging Wan 2.2 low-noise into LTX 2.3 looked promising, but the author says more layer mapping and configuration work is needed.
The post links a GitHub repository with the tested nodes and config files.
More from Multimodal
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22
- Reddit user seeks ComfyUI NSFW text-to-image and image-to-video workflows under 20 GB VRAM — hobbyist2020 · 2026-07-22
- Krea 2 users recommend a two-pass Clownshark sampler setup for sharper image details — listopalafoto · 2026-07-22
- Gemini Omni Flash turns a boat cabin into a cave in Flow by Google — chrisfirst · 2026-07-22
- A simple workflow to turn a photo into an image prompt using Gemini, Grok, or GPT Image — harshitagu72595 · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22