Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
The paper examines the convention of applying loss only to agent-authored action tokens during supervised fine-tuning (SFT) and introduces ActObs to supervise observation tokens as well. By training the model to predict action consequences without adding data or parameters, this method prevents action and observation gradients from becoming orthogonal. On Qwen3-4B, GRPO from ActObs achieves higher pass@k on Terminal-Bench 2.0 than its action-only counterpart. On Qwen3-8B, it increases pass@16 by +3.4 pp and extends advantages to cross-domain code editing on aider-polyglot with +4.2 pp at pass@1 for 4B. Joint supervision retains more entropy during RL.
On Qwen3-8B, it trades some pass@1 reliability for higher pass@k (+3.4 pp at pass@16) and solves more distinct tasks.