CO₂Jump: Google's coupled jump-process sampler fixes text-image grounding with self-correction

Upstairs_Theme2785 · reddit · 2026-09-30

A NeurIPS 2026 paper from Google, Google DeepMind and Stony Brook University tackles a mismatch in joint text-image generation: a model can describe the correct maze solution in text while drawing a different path—parallel generation alone doesn't ensure consistency.

Method: the CO₂Jump sampler uses text confidence and cross-modal attention to guide image updates during sampling; low-confidence tokens can be re-masked and regenerated, so earlier decisions can be revised. It needs one forward pass per denoising step and no extra training—the sampler works on the same fine-tuned model.

Evaluation covers image editing, maze solving and nonograms, with three new datasets (JEdit-1M, JMaze-200K, JNono-200K). Across 8–512 sampling steps, CO₂Jump was the only compared sampler that improved monotonically on both editing quality and grounding. Project page: coupled-jump.github.io

Original post →

More from Multimodal

Multimodal channel →