llama.cpp merges GLM-5.2-Vision support for local multimodal inference
QuixiAI · x · 2026-07-27
llama.cpp has merged support for GLM-5.2-Vision, adding multimodal conversion and runtime handling.
The merge:
- Registers GLM-5V with the kimivl conversion backend.
- Adds a dedicated glm5v vision projector type.
- Reuses the existing Kimi-K2.5 MoonViT3d encoder implementation.
- Updates tensor naming, the mtmd loader, graph builder, image preprocessing, dynamic token calculation, and embedding logic.
- Uses the <beginofimage> / <endofimage> delimiters.
In short, the project now has native support for this vision model in the local inference stack.
More from Infra
- Open-source profiler tracks every STT, LLM, and TTS call in self-hosted voice agents — mahimairaja · 2026-07-27
- SemiAnalysis says better memory and storage can beat a faster GPU in modern inference — rwang07 · 2026-07-27
- Building an LLM server taught one author how hard self-hosting and reliability are — MaxChamp08 · 2026-07-27
- Cheap storage makes SCD Type 2 look obsolete, says a Meta-style data engineer — Zachly · 2026-07-27
- HuggingHack adds S3, MinIO, Ollama and vLLM dispatch in a self-hosted layer — TyedalWaves · 2026-07-27
- After Copilot went unlimited, one Reddit user is deciding whether to sell a two-GPU AI rig — tweetibird · 2026-07-27