Deleting Definitions Changes Nothing: Typed Decision Models Follow Labels, Not Rules
USC · hf · 2026-10-07
The paper examines Jev-style typed decision models (used for routing, moderation, triage, outputting probabilities over options with labels and definitions): does probability follow the definitions or the labels? The authors name the phenomenon option-label bias.
Across four open-weight decision models, three ways of reading answers from a Qwen2.5 backbone, eleven tasks, and a new synthetic PolicyBench where the rule appears only in definitions, the answer is mostly the labels: deleting every definition leaves accuracy unchanged (laya-td: 0.8559 vs 0.8487) even though definitions alone support 0.7971, and renaming options to A/B raises accuracy by +0.15.
Root cause: the one unaffected system, von, differs from laya only in rendering — laya writes {label}: {definition}, von writes only the definition. Patching that single string in either direction eliminates or creates the effect (von drops 0.8511 → 0.2281 when a label contradicts its definition), locating the failure in prompt rendering rather than the constrained decision head as previously believed. A two-call test lets practitioners diagnose their own model, and four mitigations are evaluated.
More from Research
- Prediction: zeroth-order optimization methods may eventually replace backprop — j_foerst · 2026-10-07
- Berkeley's Workhorse trains humanoid G1 for whole-body manipulation from human data only — pabbeel · 2026-10-07
- ANVIL III Optimizer Claims 62% Pretraining Cost Cut at Frontier Scale, Beats Tuned Muon — kellerjordan0 · 2026-10-07
- AI safety paper highlights: reward hacking, RL debate, agent swarms — gasteigerjo · 2026-10-07
- CPT a 9B model on 2B legal tokens: how do you restore instruct and thinking? — SignificantZebra5883 · 2026-10-07
- Pathway's BDH-CQ hits 29.5% on ARC-AGI-1 at $0.0007 per task with 150M params — bendee983 · 2026-10-07