Marathoner: Ultra-Long-Horizon Autonomous Intelligence
This paper introduces Marathoner, an autonomous agentic model capable of ultra-long-horizon execution. The authors propose a comprehensive post-training pipeline featuring ultra-long-horizon task synthesis from major GitHub PRs, multi-task chaining, rejection sampling finetuning, and reinforcement learning with independent sandboxes. Additionally, a novel Later Stage Bonus Reward strategy encourages meaningful execution maneuvers during later stages, enabling the model to consistently handle tasks requiring thousands of tool calls.
Marathoner achieves consistent and substantial performance improvements over base model and even surpasses performance of strong proprietary model.
Marathoner can consistently work for 10+ hours and conduct 1000+ tool calls on highly challenging tasks.