I2T loss stays comparable across tokenizers and predicts image generation quality

peterxichen · x · 2026-09-15

From a training research thread (part 6/9): image-to-text (I2T) loss likely reflects overall joint training progress. Since every model predicts within the same text token space, I2T loss remains comparable across tokenizers, whereas text-to-image loss does not.

Key finding: I2T loss stays informative after SFT, correlating with both image generation quality and general VQA performance—making it a cross-model yardstick for multimodal training progress.

Original post →

More from Research

Research channel →