Cartesia has released Sonic-3.6, the latest version of its real-time text-to-speech model, which now holds the #1 position on both Artificial Analysis speech leaderboards. The update focuses on improved naturalness and runs on state space models rather than transformers to achieve sub-90ms time-to-first-audio latency.
- Sonic-3.6 scores 1,283 Elo on the Provider Voice board and 1,283 Elo on the Controlled Voice board.
- The model is available in beta via Cartesia's hosted API but does not offer self-hosted weights.
- Key production features include inline expression tags, instant voice cloning from 10 seconds of audio, and native alphanumerics support.
- Artificial Analysis normalizes the pricing at $49.00 per 1M characters, which is half the cost of ElevenLabs Eleven v3.
The release provides a commercially deployable streaming TTS solution that balances low latency with high naturalness for applications like customer support and media localization.