Meta's FLAT unifies image-text representations with 1D flexible-length tokens, hits 83.1 GenEval

meta · hf · 2026-09-17

Meta introduced FLAT (Flexible-Length Aligned Transmodal representations), a pre-training framework that jointly trains a shared multimodal encoder with text-to-image and image-to-text decoders, replacing the traditional two-stage setup where generative performance is bottlenecked by frozen embeddings.

Key ideas

Results

Qualitative evaluations show FLAT embeddings natively support linear interpolation, latent-space arithmetic, and zero-shot composed retrieval.

Original post →

More from Multimodal

Multimodal channel →