Build your own single-pass decision model: masking vocabulary with Qwen3-1.7B to mimic Jev

bibryam · x · 2026-10-11

Nish Tahir breaks down how "System one" decision models like Jev work and how to replicate them. Even with Structured Output, an LLM needs one forward pass per token (11 in his example) to emit valid JSON. Decision models instead assume a fixed option set (A–E), mask the vocabulary to those tokens, and pick the highest-probability answer in a single pass. The author demos this with Qwen/Qwen3-1.7B and transformers, including full code and a visualization. He also cautions that constrained output doesn't guarantee correctness, and that raw token probabilities aren't calibrated confidence scores without extra training.

Original post →

More from Research

Research channel →