ActObs: Supervising Observations During SFT Boosts Agent RL by up to 4.2pp

kastnerkyle · x · 2026-09-19

The paper "Don't Mask the Environment" introduces ActObs: instead of applying SFT loss only to agent action tokens, it also supervises the environment observation tokens already present in each trajectory, teaching the policy to predict action consequences with no extra data, parameters, sequence tokens, or forward passes.

Key results:

Mechanistic analysis: ActObs retains more entropy during RL with less policy movement, keeping the final policy closer to its SFT initialization; the authors trace the divergence to action/observation gradients becoming orthogonal early in SFT.

Related event: ActObs: Supervising Environment Tokens in SFT Boosts Agent RL(2 posts)→

Original post →

More from coding & agent

coding & agent channel →