BIT unifies text-to-image and image-to-text in one bidirectional diffusion
burny_tech · x · 2026-09-01
New research introduces BIT (Bidirectional Image-Text Diffusion Bridges), unifying text-to-image and image-to-text generation. Instead of starting from noise, the model transitions directly between text tokens and images within a single bidirectional diffusion process. Paper, demo, and code are available.
Related event: BIT Framework Unifies Text-to-Image and Image-to-Text Generation(2 posts)→
More from Multimodal
- DreamX-Creator: Native 2K Audio-Video Generation via 7B Model — GD-ML · 2026-09-01
- Generating Videos with Different Accents: Voice Consistency Issues — delveccio · 2026-09-01
- Tip: Use Gemini for first pass, then Sol for detail correction — A_K_Nain · 2026-09-01
- Japanese Sports Challenge: 30s Photorealistic Video Gen — SimplyAnnisa · 2026-09-01
- Seedance 2.5 Upgrades AI Influencer Tools with Consistency and Storytelling — aftahi_ai · 2026-09-01
- RaVayan Character Generated with Krea_ai Seedance 2.5 — CurieuxExplorer · 2026-09-01