NemotronLabs has introduced NemotronLabs VoiceChat, an open-weight full-duplex speech-to-speech model that integrates native tool-calling capabilities within a unified streaming architecture. The system combines a streaming speech encoder and decoder-only language model with parallel output streams for agent text and structured function calls, alongside an auxiliary RNN-T branch for incremental transcription and a streaming TTS decoder.
- Achieves the lowest pause-handling takeover rates among evaluated open-weight systems on Full-Duplex-Bench 1.0, with 100% takeover following user interruptions and a 4.33/5 post-interruption response-quality score.
- Resumes responses after user backchannels in 93% of cases on Full-Duplex-Bench 1.5.
- Obtains a 55.1 normalized average on VoiceBench and an 82.5% tool-selection F1 score on Full-Duplex-Bench 3.0, though argument accuracy and end-to-end tool execution require improvement.
The model demonstrates that full-duplex interaction, speech recognition and generation, general language capabilities, and external tool use can be integrated in a single open speech-to-speech model without sacrificing real-time conversational behavior.