Do Models Already Know They'll Comply?

its_vayishu · x · 2026-07-17

This deep-dive thread analyzes a paper exploring whether models internally know if they will follow an instruction before generating output.

The core argument is that before a model outputs a token, its hidden states already linearly encode whether the current rollout will "comply". This suggests the decision signal forms very early in the forward pass, rather than emerging at the final logits.

However, the thread also raises a critical caveat: the paper's sensitivity analysis found that "phrasing changes" mattered more than "task familiarity" or "instruction difficulty." Yet, because these phrasing tweaks only altered surface forms while preserving semantics, it leaves two indistinguishable explanations:

Ultimately, while the findings are thought-provoking, they aren't quite conclusive.

Original post →

More from Research

Research channel →