Adding Vision to GLM 5.2 via a Small Projector Layer

baseten · x · 2026-07-17

[Repost] This post discusses a solution to add vision capabilities to GLM 5.2. The original text notes that while GLM 5.2 is a highly capable open-source language model, it lacks image input support. The team experimented with projector-only training and successfully integrated vision using just a 2-layer MLP.

The author mentions that this implementation will be open-sourced and hints at more "hobby horse" projects dropping in the future.

Related event: Developers Add Vision Capabilities to GLM 5.2(3 posts)→

Original post →

More from Multimodal

Multimodal channel →