CO₂Jump: Google's coupled jump-process sampler fixes text-image grounding with self-correction
Upstairs_Theme2785 · reddit · 2026-09-30
A NeurIPS 2026 paper from Google, Google DeepMind and Stony Brook University tackles a mismatch in joint text-image generation: a model can describe the correct maze solution in text while drawing a different path—parallel generation alone doesn't ensure consistency.
Method: the CO₂Jump sampler uses text confidence and cross-modal attention to guide image updates during sampling; low-confidence tokens can be re-masked and regenerated, so earlier decisions can be revised. It needs one forward pass per denoising step and no extra training—the sampler works on the same fine-tuned model.
Evaluation covers image editing, maze solving and nonograms, with three new datasets (JEdit-1M, JMaze-200K, JNono-200K). Across 8–512 sampling steps, CO₂Jump was the only compared sampler that improved monotonically on both editing quality and grounding. Project page: coupled-jump.github.io
More from Multimodal
- AI painting of the day, generated with ChatGPT Image 2.5 — DeryaTR_ · 2026-10-01
- Seedance 2.5 demo brings anime-level dual-sword choreography into photorealistic cinema — SimplyAnnisa · 2026-10-01
- Midjourney --sref 3896456162 recreates 1970s Kodak film & disco aesthetics — michaelrabone · 2026-10-01
- A finished LoRA run doesn't mean it learned the style: a reproducible SDXL validation workflow — no3us · 2026-10-01
- Editor open-sources open-fusion-mcp: Claude builds editable motion graphics inside DaVinci Resolve — JohnnyLegion · 2026-10-01
- AI-generated clip: Raven Noire writing in her diary while listening to goth music — Street-Pound5762 · 2026-10-01