Tencent has made its AuK 1.5B foundation model for speech generation and editing available as open-source software, including code and model weights.

  • Supports zero-shot and instruction-based text-to-speech, content and acoustic editing, paralinguistic editing, speech enhancement, and source separation via a unified natural-language interface.
  • Includes AuK-Flash, a 4-step distilled variant for faster inference.
  • Utilizes the Qwen2.5-Omni-3B model as its MLLM encoder.
  • Released under the MIT License with weights available on Hugging Face and ModelScope.

The release allows researchers and developers to access the base model and distilled variants for various audio processing tasks.