Do Models Already Know They'll Comply?
its_vayishu · x · 2026-07-17
This deep-dive thread analyzes a paper exploring whether models internally know if they will follow an instruction before generating output.
The core argument is that before a model outputs a token, its hidden states already linearly encode whether the current rollout will "comply". This suggests the decision signal forms very early in the forward pass, rather than emerging at the final logits.
However, the thread also raises a critical caveat: the paper's sensitivity analysis found that "phrasing changes" mattered more than "task familiarity" or "instruction difficulty." Yet, because these phrasing tweaks only altered surface forms while preserving semantics, it leaves two indistinguishable explanations:
- The change simply made the constraints easier for the model to parse;
- The change genuinely altered the model's internal representation of its "commitment to comply."
Ultimately, while the findings are thought-provoking, they aren't quite conclusive.
More from Research
- Structural ensembles beat single predictions in TCR:pMHC generalization study — quaidmorris · 2026-07-22
- RSS launches under OMSF to push structural biology data modeling at scale — MoAlQuraishi · 2026-07-22
- enFoldX tops 8 neoantigen scans and an unseen-peptide benchmark — quaidmorris · 2026-07-22
- enFoldX reaches AUC 0.82 on human VDJdb and transfers to mouse at 0.76 — quaidmorris · 2026-07-22
- enFoldX gains accuracy as AF3 ensemble disagreement rises for non-binders — quaidmorris · 2026-07-22
- A 3D ray plot shows how hard this Jacobian counterexample is to read — moultano · 2026-07-22