Princeton's i1: the first fully open, reproducible text-to-image recipe, beating prior open models by 29.5 points
jiqizhixin · x · 2026-08-28
Princeton's Zhuang Liu team releases i1, a 3B-parameter DiT text-to-image model trained with flow matching, addressing the gap where every strong T2I model is either closed source or weight-only — no fully open, reproducible recipe existed before.
The key contribution is 300+ controlled experiments across 700,000 TPU v6e hours that systematically isolate what matters:
- Larger single text encoder adapters beat multiple encoders at lower cost
- Dual-stream > single-stream > cross-attention
- U-ViT skip connections consistently help
- Long captions train stronger models but hurt short-prompt performance
Built on a simple baseline with frozen text encoders and shallow fusion, i1 outperforms the previous best fully open model by 29.5 percentage points averaged across GenEval, DPG-Bench, PRISM-Bench and CVTG-2K.
More from Multimodal
- Integrating MiniMax H3 into SillyTavern turns roleplay chats into video episodes — PrinceCaspian1 · 2026-08-28
- Prompt engineering tips for MiniMax H3: Creating a handheld camera aesthetic — Jimmm90 · 2026-08-28
- Google AI Studio Eclipse demo revamped into Blood Moon using Gemini — AI_Andrew · 2026-08-28
- "How to Navigate the Singularity" trailer: IRL film featuring an AI orangutan — VoidStateKate · 2026-08-28
- Minimax H3 transformation test video demonstrates multimodal capabilities — Lutha · 2026-08-28
- Alibaba's Wan 3.0 Tops Video Edit Arena Leaderboard — arena · 2026-08-28