Alibaba's Qwen team has released Qwen-Audio-3.1, a five-model audio stack that includes the new Qwen-Audio-3.1-Realtime-plus, a full-duplex speech model designed for voice agents capable of reasoning and tool use. The release also introduces pricing reductions of up to 95% on ASR services.

  • Qwen-Audio-3.1-Realtime-plus supports 262K context tokens with function calling, web search, and structured outputs via a managed API on QwenCloud.
  • Training utilizes a three-layer approach: Think (M²-OPD), Act (executable environments with GRPO rewards), and Speak/Coordinate.
  • Benchmark improvements include Audio MultiChallenge rising from 47.12 to 52.21 and FLEURS WER falling from 9.01 to 3.98 compared to version 3.0.
  • The model reduces filler rates on Full-Duplex-Bench v3.0 from 0.7590 to 0.2960, though interruption stop latency is 1.116 seconds versus 0.383 for GPT-Realtime-2.
  • Pricing is set at $6.4 per 1M audio input tokens and $24 per 1M output tokens for text and audio.

The system aims to improve voice agent reliability by training models to decide when to speak, act, and coordinate responses within conversation history.