LLMs are General Asynchronous Agents
The authors develop an asynchronous LLM framework enabling users or agents to define inference coroutines with overlapping memory states, breaking away from traditional sequential interaction cycles. This architecture addresses non-sequential real-world use cases like voice assistants, streaming video, and monitoring systems. By evaluating Qwen 3.x models within this framework, the work demonstrates that modern LLMs can natively handle asynchronous operations for streaming video understanding, videogames, and monitoring without requiring task-specific training.
The framework lets users or agents define inference coroutines with overlapping memory states.
Qwen 3.x models are capable of asynchronous operation for streaming video understanding, videogames, and monitoring without task-specific training.