Apple Proposes Visual Concept Inference Task VICIS
Apple ML Research · rss · 2026-07-17
Apple ML Research introduced a new task called **VICIS (Visual Concept Inference from Sets)**. It aims to evaluate whether multimodal models can infer a shared concept from a set of example images and apply it to new inputs. The article notes that while existing VLMs can follow complex text instructions, they struggle with tasks requiring them to deduce concepts purely from visual context. To test this, the authors propose a setup: given a set of example images sharing a concept and a query image, the model must generate a new image that preserves the concept while aligning with the query. The paper highlights that current state-of-the-art VLMs perform poorly on this task.
More from Research
- Microsoft Research shrinks pathology models 50%+ and keeps 97% of GigaPath performance — iScienceLuvr · 2026-07-21
- OpenMHC releases 60 million hours of wearable health data for foundation models — iScienceLuvr · 2026-07-21
- Distillation alone is unlikely to explain the rise of Chinese AI models, says Reddit post — pier4r · 2026-07-21
- New papers say scaffolds explain only 1.5% of agent performance variance — gerardsans · 2026-07-21
- LLM-as-a-Coach turns judge feedback into transferable experiential knowledge — iScienceLuvr · 2026-07-21
- SWE-Pruner Pro trims coding-agent context by up to 39% using the agent’s own states — pmttyji · 2026-07-21