Community discusses definition differences between VLM and MLLM

TimDarcet · x · 2026-08-25

The post addresses the confusion surrounding the terms VLM and MLLM in academia and industry. The author notes that CLIP-style models are typically called VLMs, while LLM-based models accepting image inputs like LLaVA should be referred to as MLLMs or VLLMs. It emphasizes maintaining terminological precision in scientific communication by considering historical origins.

Original post →

More from Research

Research channel →