Do Models Already Know They'll Comply?
its_vayishu · x · 2026-07-17
This deep-dive thread analyzes a paper exploring whether models internally know if they will follow an instruction before generating output.
The core argument is that before a model outputs a token, its hidden states already linearly encode whether the current rollout will "comply". This suggests the decision signal forms very early in the forward pass, rather than emerging at the final logits.
However, the thread also raises a critical caveat: the paper's sensitivity analysis found that "phrasing changes" mattered more than "task familiarity" or "instruction difficulty." Yet, because these phrasing tweaks only altered surface forms while preserving semantics, it leaves two indistinguishable explanations:
- The change simply made the constraints easier for the model to parse;
- The change genuinely altered the model's internal representation of its "commitment to comply."
Ultimately, while the findings are thought-provoking, they aren't quite conclusive.
More from Research
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11