VLM Acts as the Eyes for GLM

MaziyarPanahi · x · 2026-07-11

GLM-5.2 was used to drive the Mol viewer for structural biology tasks, but it couldn't actually "see" what it was rendering. The author used Qwen3-VL-235B as the "eyes" to continuously check each render and provide feedback, such as noting that "the drug isn't centered and needs to be zoomed in."

This VLM-as-eyes loop iterates until the target molecule is correctly positioned in the pocket, and the final result is visualized in 3D. The author notes the entire pipeline is based on open-source models and is available on Hugging Face.

Related event: Open-source AI tackles structural biology via GLM-5.2 and Qwen3-VL self-correcting loop(5 posts)→

Original post →

More from Multimodal

Multimodal channel →