Debate on VLM Definition: Are CLIP Encoders Considered VLMs?

giffmana · x · 2026-08-26

A discussion on X involving Jitendra Malik and others debated the definition of VLM (Vision-Language Model). Some noted that CLIP image/text representation encoders are sometimes called VLMs, while others argued VLM should strictly refer to LLMs that accept image inputs.

Related event: Debate Over VLM vs MLLM Definitions(2 posts)→

Original post →

More from Models

Models channel →