Voice AI Research

Week 2026-W31
3announcements

PolyAI@polyaivoice 2 in corpus · first 2026-07-30

2026-07-30 · 2026-W31 · model · model · Dialog-RSN-1

PolyAI Dialog-RSN-1: audio-native dialog, TTS left to a separate model

PolyAI published Dialog-RSN-1, a request-based model that does turn-taking, ASR, function calling, and response, and leaves speech generation to a separate TTS.

Contact-centre stacks that need a branded TTS can take audio-aware input without moving the call onto a speech-to-speech model.

The traditional cascade loses the audio. Speech-to-speech loses control. Dialog-RSN-1 loses neither. We're introducing an entirely new approach to building voice agents, one that hears every caller's audio directly.

https://x.com/polyaivoice/status/2082819930383130879

xAI@xai 2 in corpus · first 2026-07-29

2026-07-29 · 2026-W31 · model · model · Grok Voice Think Fast 2.0

Grok Voice Think Fast 2.0: speech-to-speech at $0.08/min

xAI released Grok Voice Think Fast 2.0, a speech-to-speech model with 82.9% on AA STS, 0.70s time-to-first-audio, and $0.08/min pricing.

Full-duplex STS at $0.08/min with an A/B test on the Starlink support line. Became grok-voice-latest on 5 Aug.

OpenAI@OpenAI 3 in corpus · first 2026-07-08

2026-07-29 · 2026-W31 · model · model · GPT-Transcribe

GPT-Transcribe and GPT-Live-Transcribe: context-aware ASR

OpenAI introduced two context-aware transcription models, streaming GPT-Live-Transcribe and async GPT-Transcribe, that take prompts, keywords, and language hints.

Promptable ASR. Community cite for the Live variant: context-aware semantic accuracy 38.5% to 44.6% with free-form context.