Hearspeech-to-text7
xAIModel providermodel release2026-09-18
xAI releases Grok Voice Transcribe 2.0 speech-to-text
Puts a hyperscaler STT refresh into the same price band as Transcribe 1.0, with a named Loom deployment and a timed 1.0 sunset — peers will pressure-test the vendor benches.
Details2 sources
xAI shipped Grok Voice Transcribe 2.0, a batch and streaming STT model it says is twice as accurate as Transcribe 1.0 at the same $0.10/hr batch and $0.20/hr streaming price.
Grok Voice Transcribe 2.0
Speechmatics launches Agent STT powered by Linden 1
A dedicated agent STT surface from a long-running ASR vendor, with both Pipecat and LiveKit plugins available on day one and a public latency–accuracy claim that peers will pressure-test.
Details2 sources · 1 X post
Speechmatics released Agent STT, a speech-to-text API for production voice agents, powered by its new Linden 1 model with sub-350 ms finalisation, 55+ languages, live diarisation, and conversational events over `/v2/agent`.
Agent STT (Linden 1)
Pipecat 1.11.0 expands OpenAI, Whisper, Sarvam, AWS, and AssemblyAI STT knobs
STT latency/accuracy and vocabulary controls catch up across the major providers.
Details3 sources
pipecat-ai 1.11.0 adds OpenAIRealtimeSTT delay, OpenAISTT keywords, Whisper initial_prompt, Sarvam saaras:v4, AWS Transcribe partial_results_stability, and more AssemblyAI language_codes.
pipecat-ai 1.11.0
DeepgramModel providermodel release2026-09-17
Deepgram launches Nova-3 Pharma speech-to-text
Splits pharma entity accuracy from general medical STT — a measurable KRR claim competitors in clinical voice will be asked to match.
Details1 source · 1 X post
Deepgram released Nova-3 Pharma, a speech-to-text model trained for drug names and pharmaceutical vocabulary, available for batch and streaming on hosted and self-hosted deployments via model=nova-3-pharma.
Nova-3 Pharma
AssemblyAI launches Dictation API for finished-text STT
Bundles transcript-plus-rewrite into one billable call for push-to-talk UIs — a different request shape from streaming agent STT already in the corpus.
Details3 sources · 2 X posts
AssemblyAI released the Dictation API, a sync endpoint that returns a verbatim transcript and an LLM-cleaned finished text from one short clip on Universal-3.5 Pro, priced at $0.62 per audio hour all-in.
Dictation API
LiveKit Agents 1.8.2 adds ElevenLabs secondary languages for realtime STT
ElevenLabs multilingual realtime STT.
Details1 source
livekit-agents 1.8.2 adds secondary languages and language detection for ElevenLabs realtime STT, and migrates Speechmatics to Agent STT.
livekit-agents 1.8.2
Deepgram makes India voice endpoint generally available
Adds a managed onshore option after EU and Australia — relevant for Indian BFSI voice workloads that previously faced self-host or cross-border trade-offs.
Also Speak
Details1 source · 1 X post
Deepgram opened a generally available India regional endpoint at api.in.deepgram.com in AWS ap-south-2 (Hyderabad), with in-country storage and inference for STT, TTS, Voice Agent, and text intelligence at Global/EU/Australia pricing.
India endpoint (api.in.deepgram.com)
ThinkLLMs, speech-to-speech7
Alibaba ships Qwen3.8-LiveTranslate realtime interpretation
Vendor-claimed 2.3 s LAAL and speaker-attributed cloned speech on a 60/29-language realtime API — a direct peer for Gemini Live Translate-style interpretation workloads.
Details4 sources
Alibaba Cloud and the Qwen team released Qwen3.8-LiveTranslate, a WebSocket realtime interpretation model that understands 60 languages and speaks 29, with average lagging cut from 2.8 s to 2.3 s versus the prior generation.
Qwen3.8-LiveTranslate (qwen3.8-livetranslate-flash-realtime)
Alibaba ships Qwen-Audio 3.1 Realtime Plus for duplex voice agents
Makes 3.1 the documented default for Bailian realtime voice agents while leaving Flash as the cost-sensitive option; migration is a model ID and voice swap on the existing 3.0 protocol.
Details4 sources
Alibaba Cloud Model Studio documents qwen-audio-3.1-realtime-plus as the recommended end-to-end speech-to-speech model for voice assistants and customer-service conversations, available in Singapore and China (Beijing).
Qwen-Audio 3.1 Realtime Plus
Pipecat 1.11.0 wires Gemini 3.8 Live and extended thinking
Gemini 3.8 Live becomes selectable the same week Pipecat UI lands for builders.
Details3 sources
pipecat-ai 1.11.0 adds gemini-3.8-live and gemini-3.8-live-extended-thinking on GeminiLiveLLMService with NON_BLOCKING tools and multi-turn_complete hold.
pipecat-ai 1.11.0
StepFunModel providermodel release2026-09-15
StepFun ships StepAudio 3 Realtime topping Full-Duplex Bench at 98.9
Shanghai lab posts a duplex S2S API family that leads AA conversational dynamics while still trailing peers on first-audio latency — the quality/latency trade is the story.
Also HearSpeakOrchestrate
Details5 sources
StepFun released StepAudio 3 Realtime, a full-duplex audio-language model with Think-While-Speaking and a Voice Agent, reporting 98.9 overall on Artificial Analysis Full-Duplex Bench alongside ASR Max, TTS, Gen, and Music siblings.
StepAudio 3 Realtime
GoogleModel providermodel release2026-09-15
Google ships Gemini 3.8 Live and Extended Thinking on the Live API
A hyperscaler S2S drop with async tools + visual grounding, same day both Pipecat and LiveKit wire it — the competitive reference for GPT-Live-1-class agent stacks, with Artificial Analysis and τ-Voice numbers that need independent checking.
Details3 sources · 5 X posts
On 15 Sep Google released Gemini 3.8 Live and 3.8 Live Extended Thinking for native speech-to-speech on the Gemini Live API and AI Studio, with async tool calls, visual grounding, and mid-conversation switching across 97 languages.
Gemini 3.8 Live and 3.8 Live Extended Thinking
DeepLModel providermodel release2026-09-15
DeepL Voice adds real-time voice preservation across languages
Moves meeting translation from shared synthetic voices to per-speaker identity — a different product axis from agent STT/TTS already tracked here.
Details3 sources · 1 X post
DeepL shipped new Voice models that preserve each speaker's voice, tone, rhythm, and expression during live multilingual translation, initially across 14 languages, plus a desktop app for Zoom, Teams, and Meet.
DeepL Voice (voice preservation)
Nuance Labs raises $50M Series A for full-duplex face-to-face AI
Funds a single-model face-to-face duplex bet against cascaded avatar stacks — research preview still ahead, but NVIDIA joined the round.
Also ShowSpeak
Details3 sources
Seattle’s Nuance Labs raised a $50 million Series A led by Lightspeed, with Accel, South Park Commons, NVIDIA, and Define Ventures, to build one full-duplex audiovisual model for face-to-face conversation.
Nuance Labs foundation model
Useagents in market8
Pipecat 1.11.0 switches init clients to Pipecat UI
The default builder-facing client is now the Pipecat UI console product.
Details3 sources
pipecat-ai 1.11.0 makes pipecat init React clients use Pipecat UI (ui.pipecat.ai) instead of voice-ui-kit, with console connect, transcript, metrics, and live events.
pipecat-ai 1.11.0
MetaModel providerlaunch2026-09-17
Meta Muse beta adds outbound calls to US businesses
Consumer agent stacks now dial merchants, not just chat — product parity with Instinct Concierge (2026-09-16-instinct-concierge-calling), filed from TechCrunch plus the Muse product lead's X post rather than a Meta blog.
Also ThinkConnect
Details2 sources · 1 X post
Meta expanded the Muse beta so the agent can place outbound phone calls to US businesses, rolling first to users who had already asked Muse for calling.
Muse outbound business calling
we just expanded the @Muse beta for outbound calls to US businesses, prioritizing folks who'd asked their Muse to let us know they wanted it first :)
if you're in, give it a shot and tell us what you think. phone calling was one of our top requests and your feedback helped make it happen! 📞
https://x.com/wailord/status/2100342273854894533
Betterbot ships multifamily Voice AI with a native Leasing Dashboard CRM
Multifamily operators get phone in the same agentic context as chat and SMS on 17 September, with Entrata/Yardi/RealPage sync — a vertical peer to EliseAI-class housing voice rather than a greenfield CCaaS.
Also Connect
Details2 sources
Betterbot opened a new Voice AI lane and a Leasing Dashboard CRM so one agentic stack answers, qualifies, and routes multifamily prospect and resident calls alongside chat, SMS, and email.
Voice AI and Leasing Dashboard
Assort expands direct NextGen Enterprise EHR actions for specialty practices
Moves vertical voice agents past intake into write-back scheduling and referrals on a major ambulatory EHR — the bottleneck after the phone is answered.
Details2 sources
Assort Health announced expanded Platinum API Tier access so its AI agents can read appointment slots and create, cancel, or reschedule appointments directly in NextGen Enterprise EHR, plus referrals, notes, and payment workflows.
NextGen Enterprise EHR direct integration (expanded) · vendor NextGen Healthcare
Instinct Concierge lets the personal agent place outbound phone calls
Consumer text agents that already run email and trusted-network assistant-to-assistant contact now dial merchants for tasks with no web form — same week Meta Muse opened US business calling (2026-09-17-meta-muse-outbound-calls).
Also Connect
Details3 sources · 1 X post
Early-access Concierge has Instinct dial businesses for high-touch errands such as offline restaurant bookings, dentist cancellation lists, and cable-bill disputes.
Instinct Concierge
ElevenLabs launches Reception, an AI receptionist for small businesses
Packages ElevenAgents as an SMB receptionist SKU (calendar + after-hours) against dedicated receptionist vendors, not only API buyers; HIPAA stays on the broader platform.
Details2 sources · 1 X post
Reception by ElevenAgents answers inbound calls, books appointments from a website scan, and goes live with a phone number in minutes; plans start at $29/month with 70+ languages.
Reception (Reception.ai by ElevenAgents)
Aristotle launches nationwide voice-first AI tutoring with a $5M True Ventures seed
A consumer voice tutor sold to parents at human-tutor prices — the bet is productive struggle via live voice + whiteboard, not another homework chatbot; outcomes evidence still pending beyond engagement metrics.
Details4 sources
Aristotle opened nationwide access to its voice-first AI tutor for students roughly ages 13–18 after a closed beta of 1,000+ students and 1,500+ tutoring hours, alongside a $5 million seed led by True Ventures.
Aristotle voice tutor
Hello Patient acquires Converse Health for referral and back-office AI agents
Moves a voice-first patient-access vendor into EHR document workflows, so outbound referral calls sit on the same agent stack as inbound phones rather than a bolted-on fax tool.
Details3 sources
Hello Patient bought Converse Health so its patient-conversation agents can also do referral and fax intake, chart-driven follow-up, authorisation paperwork, and records work inside outpatient EHRs.
Converse Health acquisition (back-office agents)