Berkeley's OmniTaskonomy maps when 19 generation tasks boost 25 understanding capabilities

Berkeley · hf · 2026-09-30

Berkeley researchers ask when and how image-to-image (I2I) generation supervision improves image-to-text (I2T) understanding. Using controlled task pairs expressing the same problem in different modalities, plus OmniTaskonomy—a taxonomy of 19 I2I tasks and 25 I2T capabilities—they build a transfer map showing selective, task-dependent gains: depth prediction improves metric 3D reasoning, object pointing improves counting, jigsaw reconstruction improves 2D ordering, with surprising links like 2.5D segmentation improving category recognition. Gradient alignment correlates with transfer gains, offering a roadmap for using generation as supervision for understanding.

Original post →

More from Multimodal

Multimodal channel →