PolyAI@polyaivoice
PolyAI Dialog-RSN-1: audio-native dialog, TTS left to a separate model
PolyAI published Dialog-RSN-1, a request-based model that does turn-taking, ASR, function calling, and response, and leaves speech generation to a separate TTS.
Contact-centre stacks that need a branded TTS can take audio-aware input without moving the call onto a speech-to-speech model.
The traditional cascade loses the audio. Speech-to-speech loses control. Dialog-RSN-1 loses neither. We're introducing an entirely new approach to building voice agents, one that hears every caller's audio directly.
https://x.com/polyaivoice/status/2082819930383130879
xAI@xai
Grok Voice Think Fast 2.0: speech-to-speech at $0.08/min
xAI released Grok Voice Think Fast 2.0, a speech-to-speech model with 82.9% on AA STS, 0.70s time-to-first-audio, and $0.08/min pricing.
Full-duplex STS at $0.08/min with an A/B test on the Starlink support line. Became grok-voice-latest on 5 Aug.
OpenAI@OpenAI
GPT-Transcribe and GPT-Live-Transcribe: context-aware ASR
OpenAI introduced two context-aware transcription models, streaming GPT-Live-Transcribe and async GPT-Transcribe, that take prompts, keywords, and language hints.
Promptable ASR. Community cite for the Live variant: context-aware semantic accuracy 38.5% to 44.6% with free-form context.