CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 8 upvotes

SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL

QUESTION — How can cross-segment credit misattribution be eliminated during reinforcement learning for tool-calling LLMs?

The authors propose SLCA-GRPO, a framework resolving cross-segment credit misattribution caused by standard on-policy RL algorithms broadcasting homogeneous trajectory-level advantages to all tokens. The framework incorporates Segment-Locked Credit Assignment (SLCA), a Schema-Guided LLM Simulator (SGLS) infrastructure, and Hierarchical Rewards (HierR). By decoupling advantage estimation at the structural segment level, SLCA routes execution advantages to tool tokens and preference advantages to summary tokens without requiring intermediate rollouts. On a 7B backbone, SLCA-GRPO accelerates convergence and outperforms baseline methods across multiple benchmarks.

On a 7B backbone, SLCA-GRPO accelerates convergence and outperforms standard GRPO, ToolPO, and RLTR by +2.53 pp on in-domain evaluation.

It achieves +1.36 pp on the Berkeley Function-Calling Leaderboard (BFCL).

It achieves +9.15 pp on τ^2-Bench under the same training budgets.

YanZhanPKU · 24 Sept 2026 read the original ↗
↑