CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — evaluation 19 upvotes

Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue

QUESTION — How can the conversation initiation, silence preservation, and interruption behaviors of voice assistants be accurately evaluated in multi-party full-duplex environments?

The authors introduce Duplex-MPE, a benchmark containing 2,000 scenarios with three or four human speakers and one assistant to evaluate full-duplex speech models on when to answer, remain silent, or stop speaking. Models receive continuous audio without transcripts or turn boundaries. Evaluating five open-weight speech systems shows MiniCPM-o 4.5 leads on three scored capabilities, while a transcript-based Gemini 3.1 Pro reference responds 64.3 percentage points more often to explicit than implicit requests.

MiniCPM-o 4.5 leads on three scored capabilities, while frequent speech from other systems can coexist with inaccurate answers or failures to remain silent.

A transcript-based Gemini 3.1 Pro reference responds 64.3 percentage points more often to explicit than implicit requests; paired tests detect no significant response-rate difference for the speech systems.

ChengqianMa · 25 Sept 2026 read the original ↗
↑