The Alibaba Qwen Team has released Qwen3.8-LiveTranslate, a next-generation real-time simultaneous interpretation model that listens to live speech and returns translated text and audio while the speaker is still talking. The release introduces a new Interleave architecture designed to reduce latency and improve translation quality.

  • Average lagging (LAAL) drops from 2.8 seconds to 2.3 seconds, an approximately 18% improvement.
  • The model supports 60 languages for understanding and can speak 29 of them.
  • New capabilities include real-time speaker diarization, synchronized bilingual display, and long-context disambiguation.
  • It accepts audio and optional video inputs to assist with visual cues in noisy environments.
  • Available via WebSocket API on Alibaba Cloud Model Studio and QwenCloud.

The model is deployed as a hosted API, allowing developers to integrate real-time interpretation into applications with specific pricing tiers for audio and text inputs and outputs.