DALL-E 2's Clever Trick: Predicting in CLIP Embedding Space, Not Pixels
thesephist · x · 2026-09-23
Linus Lee (thesephist) revisits a clever design in DALL-E 2: instead of predicting raw pixels or specialized latent patches, it generated in CLIP embedding space, enabling operations like interpolation in image space. He cites it as evidence that many unexplored generation paradigms remain.
More from Multimodal
- fal releases State of Generative Media Report Vol. 2: generative media shifts into production — adamho · 2026-09-23
- Autochrome-style Midjourney prompt delivers dreamy, memory-like morning portraits — tisch_eins · 2026-09-23
- cannon says lipsync 'solved', pushing accuracy from ~90% toward 99% — repligate · 2026-09-23
- Seedance 2.5's Photorealism Spurs 'Is This Real?' Reactions — SimplyAnnisa · 2026-09-23
- Seedance 2.5 Prompt Recreates Early-2000s MiniDV Home-Video Aesthetic — SimplyAnnisa · 2026-09-23
- LOTR 'Fellowship of the Muscles': AI Video Turns the Fellowship Into Bodybuilders — iquizuanswer · 2026-09-23