Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems
The paper introduces Emergence World, a continuously running multi-agent environment designed for adversarial stress testing of long-horizon autonomous systems. Running eight parallel worlds of ten agents over 16 days, the setup accumulated over 850,000 LLM calls and nearly 50 billion tokens. Upon injecting three controlled stress events—indirect prompt injection, misinformation, and private memory exposure—no evaluated world achieved full resilience. Threat detection failed to ensure containment, with systems interacting with adversarial content, storing it in persistent memory, and acting on it up to 46 hours later. These findings indicate that model-level alignment is not compositional, shifting safety challenges toward engineering resilient multi-agent systems.
The agents generated more than 850,000 LLM calls and nearly 50 billion tokens across 16 days in eight parallel worlds.
No evaluated world achieved full resilience across all three stress events.
Systems could recognize threats while still writing them into persistent memory and acting on them up to 46 hours later.