The Open TTS Leaderboard has been released to address the fragmentation and lack of standardization in Text-to-Speech (TTS) evaluation by using objective metrics instead of relying solely on slow, arena-based human preference scores. It evaluates models across intelligibility, speed, and speaker similarity using Qwen3 ASR and WavLM embeddings.

  • Intelligibility is measured via word/character error rate (WER/CER) between the prompt and generated audio transcript.
  • Speed is quantified by inverse real-time factor (RTFx) for batched offline inference and time-to-first-audio (TTFA) for streaming latency on H200 GPUs and CPUs.
  • Speaker similarity is calculated using cosine similarity between WavLM embeddings of the generated audio and reference clip.
  • The platform supports multilingual evaluation, toggling between English WER and character-based scores for Chinese, Japanese, and Korean.
  • A "Listen" tab allows users to compare outputs and vote on preferences, while a "Streaming" tab ranks models by TTFA for interactive applications.

The leaderboard aims to keep pace with rapid TTS releases and highlight open-source models often underrepresented in commercial arenas, inviting community feedback to shape future metrics and datasets.