Hugging Face team ships 'Building VLMs' book: hands-on guide to vision-language models
andimarafioti · x · 2026-09-16
Hugging Face multimodal team members Merve Noyan, Miquel Farré, Andrés Marafioti, and Orr Zohar have released a new book, Vision-Language Models — Building VLMs with Hugging Face, available on Amazon and O'Reilly with open-source code and notebooks.
- Positioning: The authors wrote the book they wished existed when multimodal work shifted from research curiosity to engineering problem — practical guidance was scattered across blogs, docs, and hallway knowledge.
- Audience: ML engineers learn to train, fine-tune, and deploy VLMs with hands-on PyTorch and Hugging Face examples; researchers get core architectures and critical paper-reading skills; builders move from API users to system designers.
- Style: Code-first with concrete examples, using theory only to explain why things work.
Julien Chaumond amplified the announcement, and Merve Noyan called it her favorite.
More from Multimodal
- FiCA paper: instant Gaussian Codec avatars from a single portrait image — rsasaki0109 · 2026-09-16
- LLaDA-Image: 6B fully-diffusion DiT trained on 90% image-only data, no caption bottleneck — jiqizhixin · 2026-09-16
- MiniMax H3 motion graphics demo impresses, $50K challenge with Picsart opens — egeberkina · 2026-09-16
- After Suno dropped his chords, this user built a MIDI-to-.abc converter so YuE2 preserves them — Saren-WTAKO · 2026-09-16
- Hugging Face publishes illustrated guide to the 3D generation ecosystem — unofficialmerve · 2026-09-16
- Creator shares 4 imaginary tokens for Midjourney v8.2 to generate metaphorical imagery — LudovicCreator · 2026-09-16