CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 42 upvotes

Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

QUESTION — Does applying supervision to observation tokens during fine-tuning improve how agents explore under reinforcement learning?

The paper examines the convention of applying loss only to agent-authored action tokens during supervised fine-tuning (SFT) and introduces ActObs to supervise observation tokens as well. By training the model to predict action consequences without adding data or parameters, this method prevents action and observation gradients from becoming orthogonal. On Qwen3-4B, GRPO from ActObs achieves higher pass@k on Terminal-Bench 2.0 than its action-only counterpart. On Qwen3-8B, it increases pass@16 by +3.4 pp and extends advantages to cross-domain code editing on aider-polyglot with +4.2 pp at pass@1 for 4B. Joint supervision retains more entropy during RL.

On Qwen3-8B, it trades some pass@1 reliability for higher pass@k (+3.4 pp at pass@16) and solves more distinct tasks.

juzhengz · 17 Sept 2026 read the original ↗
↑