PolyAI has introduced Dialog-RSN-1, a dialog model that processes raw caller audio directly rather than relying on transcripts. It fuses turn-taking, speech recognition, function calling, and response generation into a single audio-aware system.

  • The model outputs the first token as EMPTY, ONGOING, or COMPLETE to manage turn-taking.
  • It runs as a request-based LLM probed on demand instead of an always-on stream.
  • PolyAI reports sub-300ms responses, +11% relative containment at a restaurant group, and −37% latency at an insurer.
  • The system is English-only at launch and delivered through PolyAI's platform without open weights.

Dialog-RSN-1 is already handling live production calls for existing customers, with new access available via early request.