CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — agents 192 upvotes

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

QUESTION — How can we prevent co-cheating when self-evolving search agents optimize themselves through closed-loop curriculum generation?

Self-evolving search agents suffer from co-cheating, a failure mode where the proposer and solver agree on shared errors, inflating internal rewards while external correctness stagnates. To fix this, the authors introduce CrossFit, which partitions the proposer's source documents into separate groups so that questions generated from one group are evaluated by an auxiliary solver trained on another. This decouples feedback ancestry from curriculum generation, eliminating same-source pseudo-label reproduction. Tested with Qwen models across search benchmarks, CrossFit substantially reduces false agreement and outperforms standard self-evolution and baseline search agents.

Multi-sample verification (MSV) reduces false-agreement mass from 6.1% to 5.7% and from 8.8% to 7.2%.

CrossFit reduces false agreement to 3.0% and 3.7%.

CrossFit improves average performance over standard coupled self-evolution by 8.8 and 8.4 points at 4B and 9B.

chenmeijia30 · 30 Sept 2026 read the original ↗
↑