Adobe's Chimera: Hybrid Visual Diffusion Transformer Cuts FLOPs by 7.3x

burny_tech · x · 2026-08-14

Adobe Research's paper 'Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers' proposes a hybrid DiT for long-context visual generation. It replaces most full attention with KDA linear attention, adds periodic MLA for global mixing, uses modality-aware short convs for spatiotemporal locality, and MoE for sparse capacity. The heterogeneous backbone scales predictably. Results show 7.3x fewer FLOPs than Wan 2.1 2B to reach the same loss, and zero-shot extrapolation from 5s to 30s video without length finetuning.

Original post →

More from Research

Research channel →