Adobe's Chimera: Hybrid Visual Diffusion Transformer Cuts FLOPs by 7.3x
burny_tech · x · 2026-08-14
Adobe Research's paper 'Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers' proposes a hybrid DiT for long-context visual generation. It replaces most full attention with KDA linear attention, adds periodic MLA for global mixing, uses modality-aware short convs for spatiotemporal locality, and MoE for sparse capacity. The heterogeneous backbone scales predictably. Results show 7.3x fewer FLOPs than Wan 2.1 2B to reach the same loss, and zero-shot extrapolation from 5s to 30s video without length finetuning.
More from Research
- AutoDesign: Meta-optimizing agent harnesses beats Claude Design on poster generation — dair_ai · 2026-08-15
- Qwen3.8-27B identical architecture to 3.6, gains purely from training — Course_Latter · 2026-08-15
- Notion's real-world benchmark: open models hit 94-98% quality, cost as low as $0.02 per task — ivanhzhao · 2026-08-15
- MLA+MTP May Be a Bad Combo: An Arithmetic Intensity Perspective — tokenbender · 2026-08-14
- Notion's real-world benchmark: open models hit 94-98% quality, cost as low as $0.02 per task — ivanhzhao · 2026-08-14
- OpenRouter Launches Web Search Benchmarks for Models and Configurations — AravSrinivas · 2026-08-14