The Alibaba Qwen Team has released Qwen3.8-LiveTranslate, a next-generation real-time simultaneous interpretation model that listens to live speech and returns translated text and audio while the speaker is still talking. The release introduces a new Interleave architecture designed to reduce latency and improve translation quality.
- Average lagging (LAAL) drops from 2.8 seconds to 2.3 seconds, an approximately 18% improvement.
- The model supports 60 languages for understanding and can speak 29 of them.
- New capabilities include real-time speaker diarization, synchronized bilingual display, and long-context disambiguation.
- It accepts audio and optional video inputs to assist with visual cues in noisy environments.
- Available via WebSocket API on Alibaba Cloud Model Studio and QwenCloud.
The model is deployed as a hosted API, allowing developers to integrate real-time interpretation into applications with specific pricing tiers for audio and text inputs and outputs.