AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks
The paper presents AREX-2 to advance the test-time self-improvement of LLM agents via reflection and long-horizon execution. By synthesizing long-horizon improvement trajectories from machine learning and algorithmic programming domains with verifiable feedback, the authors train an agent based on Qwen3.8-27B. The system achieves strong scores on MLE-bench Lite (81.8) and Frontier-CS (70.7), transfers to deep research benchmarks like BrowseComp (84.0), HLE (52.6), GAIA (92.2), and DeepSearchQA (93.8), and continues to improve as the budget of iteration rounds grows.
The agent achieves strong results on MLE-bench Lite (81.8) and Frontier-CS (70.7).
It transfers to deep research with 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA.
Performance keeps improving as its budget of rounds grows.