Princeton's i1: the first fully open, reproducible text-to-image recipe, beating prior open models by 29.5 points

jiqizhixin · x · 2026-08-28

Princeton's Zhuang Liu team releases i1, a 3B-parameter DiT text-to-image model trained with flow matching, addressing the gap where every strong T2I model is either closed source or weight-only — no fully open, reproducible recipe existed before.

The key contribution is 300+ controlled experiments across 700,000 TPU v6e hours that systematically isolate what matters:

Built on a simple baseline with frozen text encoders and shallow fusion, i1 outperforms the previous best fully open model by 29.5 percentage points averaged across GenEval, DPG-Bench, PRISM-Bench and CVTG-2K.

Original post →

More from Multimodal

Multimodal channel →