Amazon's ActObs: supervising observation tokens in SFT boosts agent RL exploration and pass@k

amazon · hf · 2026-09-18

Amazon researchers propose ActObs, which supervises environment observation tokens already present in agent trajectories during SFT, not just action tokens — teaching the policy to predict action consequences without extra data, parameters, or compute.

Key results:

Mechanism: action and observation gradients rapidly become orthogonal under action-only training, which degrades environment prediction below the base model; joint supervision preserves consequence prediction, keeps more entropy during RL, and better prepares the policy for exploration.

Original post →

More from Research

Research channel →