Alibaba's Qwen team has released Qwen-Audio-3.1-Realtime, a full-duplex voice model designed to handle natural conversation without waiting for a turn. The model is trained not only to understand and generate speech, but also to reason, call tools, and decide when it is appropriate to speak, making it more proactive in real-time dialogue.
According to the announcement, the model shows measurable gains in task-oriented voice interactions. On a τ-Voice adaptation, task success improves to 82.0% from 78.4%. More notably, the model's tendency to reply to background speech drops from 73% to 13%, indicating better awareness of whether speech is directed at it.
The model is available now, though the source does not specify the exact distribution channel or licensing terms. This release reflects a broader push toward voice agents that can manage complex, multi-step interactions while behaving more naturally in shared acoustic environments.