Using closed-decision Jev-like models to kill tag hallucination in dataset captioning

Iory1998 · reddit · 2026-09-21

A Reddit user proposes replacing free-form LLM captioning with closed-decision 'Jev-like' models for training dataset tagging. Instead of asking an LLM to caption openly (which hallucinates out-of-vocabulary Danbooru tags and drifts across batches), the workflow poses closed questions so the model picks from a fixed state list with probabilities — zero invented tags. A cheap first pass shortlists candidates, a Jev-like pass confirms them, and the frozen tag list is then handed to an LLM purely as a translator to produce natural-language descriptions, yielding both tag strings and prose in one run. The author hasn't built it yet and invites feedback on where it breaks.

Related event: Reddit Proposal: Closed Decision Models to Fix Dataset Captioning Hallucinations(2 posts)→

Original post →

More from Multimodal

Multimodal channel →