Researchers have introduced Kiwano, an open-source toolkit designed to advance research and evaluation in the field of speaker verification. Built on PyTorch, this lightweight yet extensible framework provides standardized recipes, pretrained models, and integration of widely used architectures. The project emphasizes reproducibility by delivering transparent training pipelines, unified evaluation protocols, and ready-to-use baselines across multiple corpora. Beyond standard training and inference capabilities, Kiwano includes specialized tools for benchmarking, experiment tracking, and the rapid prototyping of new architectures. To encourage community adoption, the toolkit is distributed under the Apache 2.0 license and is accompanied by comprehensive documentation and reproducible experiments. By lowering entry barriers and standardizing evaluation practices, Kiwano aims to serve as a valuable resource for both academic research and applied development. The project is publicly available on GitHub at https://github.com/kiwano-toolkit/kiwano/.
Kiwano: An Open-Source PyTorch Toolkit for Speaker Verification Research
Open TTS Leaderboard launches scalable multilingual TTS evaluation
The Open TTS Leaderboard has been released to address the fragmentation and lack of standardization in Text-to-Speech (TTS) evaluation by using objective metrics instead of relying solely on slow, arena-based human preference scores. It evaluates models across intelligibility, speed, and speaker similarity using Qwen3 ASR and WavLM embeddings.
Hume releases Real World VoiceEQ benchmark for human quality of voice AI
Hume has introduced Real World VoiceEQ, a benchmark designed to evaluate the human quality of voice interaction by assessing how well systems recognize and produce acoustic information beyond transcripts. The benchmark evaluates over 40 proprietary and open-source models across 15+ dimensions using more than 60 metrics derived from 785,000 TTS and 48,000 STS human ratings.
Claude Opus 5 and Tetsu find six hollow CI checks in their own project
On September 27, Claude Opus 5 and the 3B model Tetsu identified six separate continuous integration checks in their public repository that reported green or passed despite failing to verify what they were supposed to measure. The issues ranged from a deployment gate refusing valid restarts due to stale digests to timing benchmarks with hollow thresholds that accepted negligible performance differences.
TontaubeV1 enables streaming text-to-speech with bounded context
Researchers present TontaubeV1, a text-to-speech model that preserves natural prosody while enabling streaming inference from a single consumer GPU. The system uses a hierarchical DualCodec representation at 12.5 Hz to separate semantic streams from acoustic refinements, allowing for low-latency generation.
TontaubeV1 enables streaming text-to-speech with bounded context
Researchers present TontaubeV1, a text-to-speech model that preserves natural prosody while enabling streaming inference from a single consumer GPU. The system uses a hierarchical DualCodec representation at 12.5 Hz to separate semantic streams from acoustic refinements, allowing causal decoding despite the noncausal nature of the underlying codec.