CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — agents 34 upvotes

One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents

QUESTION — How can we overcome the category see-saw effect during reinforcement learning for software engineering agents?

This paper proposes a category-aware expert-training and policy-integration framework to address uneven progress across task categories during pooled agentic reinforcement learning for software engineering. By combining Agentic-miniRL, the Refresh-Repair-Expand (RRE) strategy, and multi-axis labeling via SWE Labeler, category-specific experts are trained. A label-routed multi-teacher on-policy distillation (MOPD) method with ReLU-gated reward extrapolation then consolidates these experts into a single deployable student without requiring external models. The final policy achieves a mean resolution of 58.04% on Pro-618 and 59.00% on SWE-bench Multilingual, improving over the base model.

The final MOPD policy achieves mean resolution of 58.04% on Pro-618 and 59.00% on SWE-bench Multilingual.

It improves over the base model by 5.39 and 2.78 percentage points, respectively.

Williams07 · 20 Sept 2026 read the original ↗
↑