IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis
QUESTION — How can role coupling and context noise be mitigated in deep search agents through iterative synthesis?
The study proposes IterSynth, a deep search agent architecture that decouples planning from evidence synthesis and uses an evolving summary state to reduce context noise. To train this paradigm, the authors introduce Role-Decoupled Policy Optimization (RDPO), combining terminal outcome rewards with turn-level rubric evaluations for precise credit assignment. Experiments across five long-horizon deep-search benchmarks demonstrate that IterSynth-8B achieves an average score of 50.7, surpassing the strongest prior leq8B agent by +4.2%.
IterSynth-8B achieves an average score of 50.7, surpassing the strongest prior leq8B agent by +4.2%.
wuxingyu · 24 Sept 2026
read the original ↗