False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
Self-evolving search agents suffer from co-cheating, a failure mode where the proposer and solver agree on shared errors, inflating internal rewards while external correctness stagnates. To fix this, the authors introduce CrossFit, which partitions the proposer's source documents into separate groups so that questions generated from one group are evaluated by an auxiliary solver trained on another. This decouples feedback ancestry from curriculum generation, eliminating same-source pseudo-label reproduction. Tested with Qwen models across search benchmarks, CrossFit substantially reduces false agreement and outperforms standard self-evolution and baseline search agents.
Multi-sample verification (MSV) reduces false-agreement mass from 6.1% to 5.7% and from 8.8% to 7.2%.
CrossFit reduces false agreement to 3.0% and 3.7%.
CrossFit improves average performance over standard coupled self-evolution by 8.8 and 8.4 points at 4B and 9B.