A developer has open-sourced Speakrail, a fully-local, full-duplex voice agent that runs on a single RTX 4090 and rivals GPT-Live on certain benchmarks. The system combines Voxtral Realtime for speech-to-text, a microturn-finetuned Gemma 4 12B model, and Breeze TTS 2 to enable low-latency interactions with interruptions and backchannels.

  • The pipeline uses a turn-taking head attached to Voxtral Realtime and speculative LLM+TTS firing to minimize latency.
  • The core language model is Gemma 4 12B QAT, finetuned on synthetic data with microturn control tokens like <interject> and <listen>.
  • A "think while talking" mechanism allows a base Gemma 4 12B int4 to write thinking notes for complex reasoning tasks.
  • The system supports quiet mode and interrupt mode, allowing the model to listen or speak based on explicit user control.

The release provides a privacy-preserving alternative to cloud-based voice assistants for users with sufficient local hardware, though it currently has limitations regarding context length and commercial licensing of the TTS component.