Alibaba / Qwen@alibaba_cloud
Alibaba ships Qwen3.8-LiveTranslate realtime interpretation
Alibaba Cloud and the Qwen team released Qwen3.8-LiveTranslate, a WebSocket realtime interpretation model that understands 60 languages and speaks 29, with average lagging cut from 2.8 s to 2.3 s versus the prior generation.
Puts a dated 2.3 s LAAL claim and speaker-attributed cloned speech on a 60/29-language realtime API — a direct peer for Gemini Live Translate-style interpretation workloads.
Alibaba ships Qwen-Audio 3.1 Realtime Plus for duplex voice agents
Alibaba Cloud Model Studio documents qwen-audio-3.1-realtime-plus as the recommended end-to-end speech-to-speech model for voice assistants and customer-service conversations, available in Singapore and China (Beijing).
Makes 3.1 the documented default for Bailian realtime voice agents while leaving Flash as the cost-sensitive option; migration is a model ID and voice swap on the existing 3.0 protocol.
xAI@xai
xAI releases Grok Voice Transcribe 2.0 speech-to-text
xAI shipped Grok Voice Transcribe 2.0, a batch and streaming STT model it says is twice as accurate as Transcribe 1.0 at the same $0.10/hr batch and $0.20/hr streaming price.
Puts a hyperscaler STT refresh into the same price band as Transcribe 1.0, with a named Loom deployment and a timed 1.0 sunset — peers will pressure-test the vendor benches.
Speechmatics@Speechmatics
Speechmatics launches Agent STT powered by Linden 1
Speechmatics released Agent STT, a speech-to-text API for production voice agents, powered by its new Linden 1 model with sub-350 ms finalisation, 55+ languages, live diarisation, and conversational events over `/v2/agent`.
A dedicated agent STT surface from a long-running ASR vendor, with Pipecat/LiveKit plugins on day one and a public latency–accuracy claim that peers will pressure-test.
That’s why we built Agent STT.
https://x.com/Speechmatics/status/2100607821846921567
Hume AI@hume_ai
Hume publishes Voice Controllability Leaderboard for TTS direction-following
Hume published a Voice Controllability Leaderboard scoring 17 TTS models on voice design, instruction-following, inline tags, and role fit with blind human raters across 13 languages.
Separates identity cloning from controllability. Same model can win motivational casting and lose meditation; Fish's voice-design swings are the clearest example in the board.
A result from the Voice Controllability leaderboard we published today: Fish's voice design model won nearly every head to head for motivational speaker voices and lost nearly every one for meditation voices. Same model, same prompt format, and the job you're casting for flipped
https://x.com/hume_ai/status/2100584836864024785
Deepgram@DeepgramAI
Deepgram launches Nova-3 Pharma speech-to-text
Deepgram released Nova-3 Pharma, a speech-to-text model trained for drug names and pharmaceutical vocabulary, available for batch and streaming on hosted and self-hosted deployments via model=nova-3-pharma.
Splits pharma entity accuracy from general medical STT — a measurable KRR claim competitors in clinical voice will be asked to match.
Introducing Nova-3 Pharma, the first speech-to-text model purpose-built for the pharmaceutical industry.
https://x.com/DeepgramAI/status/2100614929573093694
Deepgram makes India voice endpoint generally available
Deepgram opened a generally available India regional endpoint at api.in.deepgram.com in AWS ap-south-2 (Hyderabad), with in-country storage and inference for STT, TTS, Voice Agent, and text intelligence at Global/EU/Australia pricing.
Adds a managed onshore option after EU and Australia — relevant for Indian BFSI voice workloads that previously faced self-host or cross-border trade-offs.
The Deepgram India endpoint is now generally available.
https://x.com/DeepgramAI/status/2099894045023641666
Cartesia@cartesia
Cartesia ships Multilingual Voices for Sonic TTS
Cartesia launched Multilingual Voices so one cloned or library voice can speak multiple languages and accents natively while keeping the same voice ID, demonstrated with SF bakery Baklavastory.
Moves brand-voice localisation from one-voice-per-locale cloning to accent adds on a single ID — useful for agents that already run Sonic-3.6.
Introducing Multilingual Voices:
https://x.com/cartesia/status/2100630777407091096
https://x.com/krandiash/status/2100630035208323305
Assort Health@assort_health
Assort expands direct NextGen Enterprise EHR actions for specialty practices
Assort Health announced expanded Platinum API Tier access so its AI agents can read appointment slots and create, cancel, or reschedule appointments directly in NextGen Enterprise EHR, plus referrals, notes, and payment workflows.
Moves vertical voice agents past intake into write-back scheduling and referrals on a major ambulatory EHR — the bottleneck after the phone is answered.
Tavus@tavus
Tavus ships Memories for long-term PAL relationships
Memories gives each PAL–person pair a Profile and Timeline that update after calls, so returning conversations continue with prior goals, events, and preferences without re-explaining.
After Phoenix-4.5, CVI gets relationship continuity: post-call consolidation plus inspectable/editable stores, with Pinned Memories for handoff context.
SoundHound AI@SoundHound
SoundHound ships Human Assisted Resolution inside OASYS
SoundHound launched Human Assisted Resolution (HAR), an opt-in OASYS feature that lets an AI agent ask a colleague a specific question in real time and finish the conversation without a full handoff.
Separates brief human judgment from escalation — a concrete HITL pattern for branded voice/chat agents that otherwise double-pay for transfer.
Today we're launching Human Assisted Resolution (HAR).
https://x.com/SoundHound/status/2100253157674582423
ElevenLabs@ElevenLabs
ElevenLabs launches Reception, an AI receptionist for small businesses
Reception by ElevenAgents answers inbound calls, books appointments from a website scan, and goes live with a phone number in minutes; plans start at $29/month with 70+ languages.
Packages ElevenAgents as an SMB receptionist SKU (calendar + after-hours) against dedicated receptionist vendors, not only API buyers; HIPAA stays on the broader platform.
AssemblyAI@AssemblyAI
AssemblyAI launches Dictation API for finished-text STT
AssemblyAI released the Dictation API, a sync endpoint that returns a verbatim transcript and an LLM-cleaned finished text from one short clip on Universal-3.5 Pro, priced at $0.62 per audio hour all-in.
Bundles transcript-plus-rewrite into one billable call for push-to-talk UIs — a different request shape from streaming agent STT already in the corpus.
Dear developers, Forms are broken.
https://x.com/AssemblyAI/status/2100710329282101434
Blurt: push-to-talk dictation on the AssemblyAI Dictation API.
https://x.com/AssemblyAI/status/2100971811584516172
Hello Patient@HelloPatient
Hello Patient acquires Converse Health for referral and back-office AI agents
Hello Patient bought Converse Health so its patient-conversation agents can also do referral and fax intake, chart-driven follow-up, authorisation paperwork, and records work inside outpatient EHRs.
Moves a voice-first patient-access vendor into EHR document workflows, so outbound referral calls sit on the same agent stack as inbound phones rather than a bolted-on fax tool.
Google@Google
Google ships Gemini 3.8 Live and Extended Thinking on the Live API
On 15 Sep Google released Gemini 3.8 Live and 3.8 Live Extended Thinking for native speech-to-speech on the Gemini Live API and AI Studio, with async tool calls, visual grounding, and mid-conversation switching across 97 languages.
A hyperscaler S2S drop with async tools + visual grounding, same day LiveKit/Pipecat wire it — the competitive reference for GPT-Live-1-class agent stacks, with AA / τ-Voice numbers that need independent checking.
DeepL@DeepLcom
DeepL Voice adds real-time voice preservation across languages
DeepL shipped new Voice models that preserve each speaker's voice, tone, rhythm, and expression during live multilingual translation, initially across 14 languages, plus a desktop app for Zoom, Teams, and Meet.
Moves meeting translation from shared synthetic voices to per-speaker identity — a different product axis from agent STT/TTS already tracked here.
Introducing voice preservation in DeepL Voice.
https://x.com/DeepLcom/status/2099831527404154912