Using closed-decision Jev-like models to kill tag hallucination in dataset captioning
Iory1998 · reddit · 2026-09-21
A Reddit user proposes replacing free-form LLM captioning with closed-decision 'Jev-like' models for training dataset tagging. Instead of asking an LLM to caption openly (which hallucinates out-of-vocabulary Danbooru tags and drifts across batches), the workflow poses closed questions so the model picks from a fixed state list with probabilities — zero invented tags. A cheap first pass shortlists candidates, a Jev-like pass confirms them, and the frozen tag list is then handed to an LLM purely as a translator to produce natural-language descriptions, yielding both tag strings and prose in one run. The author hasn't built it yet and invites feedback on where it breaks.
More from Multimodal
- Uncensored Qwen-Image-2.1 GGUF quantization trends on Hugging Face — abenzerps · 2026-09-21
- Custom CUDA shim runs Stable Diffusion on Mac faster than RTX 5090 on Windows, up to 61% quicker — LioDavinchy · 2026-09-21
- Qwen-Image-2.1 hands-on: trails Krea2 in T2I quality and Flux2Klein in editing flexibility — Capitan01R- · 2026-09-21
- Qwen 2.1 T2I & edit test: Reddit calls it the open-source Nano Banana moment — LongjumpingGur7623 · 2026-09-21
- GPT-Image 2.5 proves surprisingly good at pixel art, animated via Seedance 2.5 — Aiden_Tech_Ai · 2026-09-21
- Qwen-Image 2.1 runs on SGLang: image generation in 18.7s on a single RTX 4090 — Alibaba_Qwen · 2026-09-21