Realtime-Venus: A full-duplex interaction system with asynchronous delegation
The authors present Realtime-Venus, a proactive full-duplex interaction system comprising two separately trained 9B models: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction. The system features a dual-loop runtime that coordinates live frontend interaction with background reasoning and tool execution through an asynchronous harness. Realtime-Venus-Omni achieves the highest scores on six of eight video benchmarks (StreamingBench at 70.2%, OVO-Bench at 64.7%, and Daily-Omni at 81.3%), while Realtime-Venus-Audio leads compared models across audio benchmarks such as MMAU (78.0%), MMAU-Pro (63.2%), Llama Questions (83.8%), and Speech CMMLU (67.8%), and responds to 75% of user interruptions on Full-Duplex-Bench v1.5.
Realtime-Venus-Omni achieves the highest scores on six of eight video benchmarks, including StreamingBench (70.2%), OVO-Bench (64.7%), and Daily-Omni (81.3%).
Realtime-Venus-Audio leads compared models on MMAU (78.0%), MMAU-Pro (63.2%), Llama Questions (83.8%), and Speech CMMLU (67.8%), while matching the best VoiceBench AlpacaEval score of 4.81.
On Full-Duplex-Bench v1.5, Realtime-Venus-Audio responds to 75% of user interruptions and achieves continuation rates of 97%, 88%, and 86% under backchannels, other-directed speech, and background speech, respectively.