Apple Proposes Visual Concept Inference Task VICIS

Apple ML Research · rss · 2026-07-17

Apple ML Research introduced a new task called **VICIS (Visual Concept Inference from Sets)**. It aims to evaluate whether multimodal models can infer a shared concept from a set of example images and apply it to new inputs. The article notes that while existing VLMs can follow complex text instructions, they struggle with tasks requiring them to deduce concepts purely from visual context. To test this, the authors propose a setup: given a set of example images sharing a concept and a query image, the model must generate a new image that preserves the concept while aligning with the query. The paper highlights that current state-of-the-art VLMs perform poorly on this task.

Original post →

More from Research

Research channel →