FireRedTeam has released two new open-source audio models: FireRedAudio, a general-purpose audio language model, and FireRedTTS3, a unified speech generation and editing system.
- FireRedAudio is built on a shared 9B-parameter LLM with decoupled continuous representations, supporting ASR, audio understanding, zero-shot TTS, instruct TTS, semantic/acoustic speech editing, and temporal grounding for recordings up to one hour long.
- FireRedTTS3-Base enables zero-shot voice cloning across 24 languages and 21 Chinese dialects, achieving competitive results on MiniMax-MLS-Test and Seed-TTS-eval benchmarks.
- FireRedTTS3-Instruct allows for natural-language voice design and free-form speech editing (semantic and acoustic) within a single unified model.
These releases provide users with comprehensive tools for both understanding complex audio contexts and generating high-fidelity, multilingual speech.