StepAudio 3 Realtime Technical Report
The paper presents StepAudio 3 Realtime, an audio-language foundation model centered on a continuous listen-converse-think-act loop. To resolve the tension between deep deliberation and latency, it introduces Think-While-Speaking, performing private reasoning in parallel with spoken delivery. Additionally, an integrated voice agent manages asynchronous tool execution smoothly. StepAudio 3 Realtime achieves strong metrics, including a 73.0 macro average on StepAudioChat, 90.6 on the MMSU benchmark, 98.9 Overall on the Artificial Analysis Full-Duplex Bench, and a 56.0% macro task-success rate on τ-Voice.
In reasoning mode, StepAudio 3 reaches a 73.0 macro average on StepAudioChat.
It achieves an exceptional 90.6 on the MMSU benchmark and 98.9 Overall on the Artificial Analysis Full-Duplex Bench.
It attains a 56.0% macro task-success rate on τ-Voice.