Extending Jev Mode to Images: Constrained llama.cpp Outputs as Image Selections

opUserZero · reddit · 2026-09-29

Codacus built Jev mode (constrained decoding) for llama.cpp; the author extended the idea to images with an agent plus harness — the constrained answer is directly an image selection, skipping the decode step entirely: no captioning pause, just a decision based on one image or a group.

Use cases: ask the same question over a batch of images for classification, or hand the model 20 images and have it pick the one containing a rubber duck. PR and a YouTube explainer (made with Codacus's RenderDiv framework) are linked.

Original post →

More from coding & agent

coding & agent channel →