Deleting Definitions Changes Nothing: Typed Decision Models Follow Labels, Not Rules

USC · hf · 2026-10-07

The paper examines Jev-style typed decision models (used for routing, moderation, triage, outputting probabilities over options with labels and definitions): does probability follow the definitions or the labels? The authors name the phenomenon option-label bias.

Across four open-weight decision models, three ways of reading answers from a Qwen2.5 backbone, eleven tasks, and a new synthetic PolicyBench where the rule appears only in definitions, the answer is mostly the labels: deleting every definition leaves accuracy unchanged (laya-td: 0.8559 vs 0.8487) even though definitions alone support 0.7971, and renaming options to A/B raises accuracy by +0.15.

Root cause: the one unaffected system, von, differs from laya only in rendering — laya writes {label}: {definition}, von writes only the definition. Patching that single string in either direction eliminates or creates the effect (von drops 0.8511 → 0.2281 when a label contradicts its definition), locating the failure in prompt rendering rather than the constrained decision head as previously believed. A two-call test lets practitioners diagnose their own model, and four mitigations are evaluated.

Original post →

More from Research

Research channel →