Users look for uncensored VLMs that can caption explicit images accurately
TekeshiX · reddit · 2026-08-04
A Reddit user asks which LLM/VLM is best for accurately describing images, including uncensored or NSFW content.
- They previously used Gemini with a jailbreak prompt for captioning explicit images for later WAN 2.2 video generation.
- Gemini now blocks those images, so they are looking for alternatives.
- Models mentioned include Qwen3.5-VL, InternVL, Gemma 4, and JoyCaption, with the user asking whether uncensored variants are needed.
- The request also specifies a 16–48 GB VRAM budget and says standalone or ComfyUI use is fine.
Related event: Developers Seek Uncensored VLMs for Explicit Image Captioning(2 posts)→
More from Models
- Qwen3-VL-32B variant starts trending on Hugging Face — ethanfel · 2026-08-04
- Dev Vows to Turn US AI Bear If DeepSeek Nails Complex Task — yacineMTB · 2026-08-04
- DeepSeek Flash takes a reverse-engineering task ChatGPT refused — yacineMTB · 2026-08-04
- A model that ‘won 10 Fields Medals’ still recommended a cafe closed two years ago — bosmeny · 2026-08-04
- Open models still have one big advantage: you can read the reasoning thread — jamesdouma · 2026-08-04
- opencode’s model picker now includes a Canadian model — yacineMTB · 2026-08-04