SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL
The authors propose SLCA-GRPO, a framework resolving cross-segment credit misattribution caused by standard on-policy RL algorithms broadcasting homogeneous trajectory-level advantages to all tokens. The framework incorporates Segment-Locked Credit Assignment (SLCA), a Schema-Guided LLM Simulator (SGLS) infrastructure, and Hierarchical Rewards (HierR). By decoupling advantage estimation at the structural segment level, SLCA routes execution advantages to tool tokens and preference advantages to summary tokens without requiring intermediate rollouts. On a 7B backbone, SLCA-GRPO accelerates convergence and outperforms baseline methods across multiple benchmarks.
On a 7B backbone, SLCA-GRPO accelerates convergence and outperforms standard GRPO, ToolPO, and RLTR by +2.53 pp on in-domain evaluation.
It achieves +1.36 pp on the Berkeley Function-Calling Leaderboard (BFCL).
It achieves +9.15 pp on τ^2-Bench under the same training budgets.