Self-Organizing Agent Teams Learn to Reason Together
The study introduces Self-Organizing Agent Teams (SAT), a framework where fixed teams of AI agents learn reusable collaboration strategies from past interactions to organize roles, conversational phases, and information flow. Instead of fixed protocols, these teams perform collaborative computation by exchanging, challenging, and synthesizing partial reasoning. Evaluated across five mathematics and physics benchmarks, self-organizing teams achieve an average 66.7% accuracy, compared to 48.8% for their strongest individual member, and exceed a perfect router by 13.4 points on AIME 2026.
Across five mathematics and physics benchmarks, self-organizing teams average 66.7% accuracy, versus 48.8% for their strongest member.
On AIME 2026, they exceed this router by 13.4 points.
Demonstrability strongly tracks improvement over the strongest member with Spearman ρ=0.90, p=0.005.