hotwirednews Install bot

Week 37

7–13 Sep 20262026-W37 · 25 filings · 18 companies · 9 net-new

The Speak layer led the week with 6 of 25 filings, from Daily / Pipecat, Twilio, HeyGen and Gradium.

Hearspeech-to-text3

Daily / PipecatPlatformlaunch2026-09-11

Pipecat 1.10.0 moves Speechmatics to Agent STT across 1.9.0 and 1.10.0

Same-week Pipecat 1.9.0–1.10.0 hear minors collapse to one card; Speechmatics Agent STT migration stays the lead with Meta and AssemblyAI Sync folded in.

Details4 sources

pipecat-ai 1.10.0 retargets SpeechmaticsSTTService to Agent STT (linden-1) with reconnect backoff and Gradium turn detection; the same ISO-week hear cluster folds 1.9.0 MetaSTTService, AssemblyAISyncSTTService, and eager-EOT frames.

pipecat-ai 1.10.0 (cluster with 1.9.0)

LiveKitPlatformlaunch2026-09-10

LiveKit Agents 1.8.1 adds Meta Muse Voice streaming STT

Muse STT plugin for Agents.

Details1 source

livekit-agents 1.8.1 adds a Meta Muse Voice streaming STT plugin and Speechmatics end_of_turn_config/vad_config options.

livekit-agents 1.8.1

Blue Machines AIModel providermodel release2026-09-07net-new

Blue Machines AI launches Aurora STT for Indian BFSI calls

India BFSI code-mixed phone STT with entity-level scoring — a vertical hear model against general ASR, still vendor-internal benches.

Also Use

Details4 sources

Blue Machines AI launched Aurora, a multilingual speech-to-text model built for code-mixed Indian banking and insurance phone conversations, with internal semantic WER of 1.51% English and 2.43% Hindi BFSI.

Aurora

Speaktext-to-speech6

Daily / PipecatPlatformAPI2026-09-11

Pipecat 1.10.0 adds Smallest continuation API across 1.9.0 and 1.10.0

Same-week Pipecat 1.9.0–1.10.0 speak minors collapse to one card; Smallest continuation and LMNT EOL lead, with the 1.9.0 settings wave folded in.

Details4 sources

pipecat-ai 1.10.0 switches SmallestTTSService to the continuation API with shared context_id per LLM turn and removes LmntTTSService; the same ISO-week speak cluster folds 1.9.0 Azure, Hume, Deepgram, Soniox, and Sarvam TTS settings.

pipecat-ai 1.10.0 (cluster with 1.9.0)

TwilioPlatformlaunch2026-09-09

Twilio opens ElevenLabs voices for Say in public beta

Puts low-latency ElevenLabs TTS on stock Twilio IVR and announce paths without a custom Media Streams bridge — a different surface from Agent Connect's LLM runtime.

Also Connect

Details1 source

Twilio made ElevenLabs Flash v2 and Flash v2.5 voices available in public beta on the Say verb via voice=ElevenLabs.<VoiceID> in TwiML or TwiML Bins.

ElevenLabs Voices for Say · vendor ElevenLabs

HeyGenPlatformlaunch2026-09-09net-new

HeyGen Professional Voice Clone trains a dedicated voice from 20+ minutes of audio via API

Avatar video quality has outrun voice fidelity; a studio-grade clone behind the same API as LiveAvatar closes that gap for builders who already have 20 minutes of clean speaker audio.

Also Show

Details2 sources · 2 X posts

HeyGen opened Professional Voice Clone on the API on 9 Sep: a paid per-slot HeyGen Voice adapter trained on 1–10 recordings totaling at least 20 minutes of one speaker.

Professional Voice Clone

GradiumModel providerpartnership2026-09-09

Gradium TTS lands on LiveKit Inference free until 9 Oct

Removes a second vendor account for LiveKit shops that want Gradium voices — promo pricing ends 9 October 2026, and custom/community Gradium voices still need the plugin path.

Details4 sources · 2 X posts

Gradium TTS is selectable inside LiveKit Inference as gradium/default, so LiveKit Agents can use Gradium voices without a separate Gradium API key.

Gradium TTS on LiveKit Inference · vendor LiveKit

Gradium TTS is now available on @livekit Inference. First month free, so every developer can test it.

https://x.com/GradiumAI/status/2097753691088392385

Gradium TTS is now available on LiveKit Inference, free until Oct. 9.

https://x.com/livekit/status/2097718917401653471
GradiumModel providermodel release2026-09-09

Gradium retrains default TTS after Coval starts scoring leading silence in TTFA

Labs that game time-to-first-byte with silent frames get caught when evals score what a caller hears; Gradium’s response shows the metric now steers shipping.

Details2 sources · 2 X posts

After Coval’s June perceived-TTFA change exposed ~225 ms median leading silence in Gradium’s stream, Gradium shipped a new default that emits audible audio in the first 10 ms frame (median TTFA ~214 ms).

Gradium default TTS (perceived TTFA retrain)

GradiumModel providerlaunch2026-09-07

Gradium launches Voice Design for prompt-to-voice synthesis

Puts prompt-built voices on the same latency path as Gradium catalog TTS, so agent teams can match regional accents without cloning a real speaker — win rates are vendor-run accent tests, not telephony ASR benches.

Details5 sources · 2 X posts

Gradium shipped Voice Design, which turns a written speaker description into up to five synthetic voice candidates that promote to permanent voice_ids on the same TTS endpoints as catalog voices.

Voice Design

Today, we're launching Voice Design. Write a prompt, create a voice.

https://x.com/GradiumAI/status/2097359468963004826

Native speakers rated 7,000+ blind comparisons and put Gradium first for accent matching.

https://x.com/GradiumAI/status/2097716535183777976

Showavatars, video1

TavusModel providermodel release2026-09-10

Tavus Phoenix-4.5 renders full upper-body PALs at 134 ms audio-to-video

Realtime avatar stacks that only animate a face leave body motion static; whole-frame generation plus sub-150 ms A2V is the bar conversational video agents now have to clear.

Details2 sources · 1 X post

Tavus shipped Phoenix-4.5 on 10 Sep: a full-frame generative renderer that moves head, shoulders, posture, and torso with speech, at 134 ms audio-to-video.

Phoenix-4.5

ThinkLLMs, speech-to-speech4

Daily / PipecatPlatformlaunch2026-09-10

Pipecat 1.9.0 adds OpenAI Live speech-to-speech service

Pipecat and LiveKit both offered a first-class OpenAI Live path on 10 Sep 2026.

Details3 sources

pipecat-ai 1.9.0 adds OpenAILiveLLMService for gpt-live-1 with ResponsesDelegation and ClientDelegation via BackendLLMWorker.

pipecat-ai 1.9.0

OpenAIModel providerAPI2026-09-10

GPT-Live-1 in the API: full-duplex voice at $0.05/min

Community cite: ~0.798s turn-taking versus 1.41s for GPT-Realtime-2.1. Backend model billed separately.

Details3 sources

OpenAI brought ChatGPT’s full-duplex GPT-Live-1 voice model to the developer API at $0.05/min for the voice layer, with backend models billed separately.

GPT-Live-1

LiveKitPlatformlaunch2026-09-10

LiveKit Agents 1.8.1 adds DuplexModel for GPT-Live

Full-duplex speech models become a first-class Agents session type.

Details1 source

livekit-agents 1.8.1 introduces DuplexModel so a session can run a full-duplex speech model, with GPTLiveModel as the first implementation.

livekit-agents 1.8.1

TencentModel providermodel release2026-09-08net-new

Tencent Hunyuan releases Gander duplex omni interaction model

Open weights for a GPT-Live-style duplex-plus-agent split — turn-taking leads the cited bench, while end-to-end tool Pass@1 still trails closed realtime systems and vision regresses versus the MiniCPM-o base.

Details3 sources

Tencent's Hunyuan Speech team released Gander, an open ~9B duplex model that streams speech and vision while a plug-in Brain runs longer agent tasks asynchronously.

Gander (Omni Interaction Agent)

Orchestrateframeworks, evals5

SierraPlatformlaunch2026-09-10

Sierra ships multimodal agents mixing voice, text and visuals

Treats visuals as part of the voice-agent turn, not a hand-off to a separate app — the post is feature narrative without named customer go-lives or latency numbers.

Details2 sources · 1 X post

Sierra's multimodal agents keep voice, text and on-screen visuals in one conversation so customers can talk, view options and tap without restarting the session.

Multimodal agents

Our multimodal agents bring voice, text, and visuals into the same conversation.

https://x.com/SierraPlatform/status/2098145890842382623
Daily / PipecatPlatformlaunch2026-09-10

Pipecat 1.9.0 adds FlowConfig and eval simulations

YAML flows plus simulated callers make agent graphs and release gates share one release.

Details3 sources

pipecat-ai 1.9.0 adds FlowConfig for data-driven Pipecat Flows, LatencyBreakdown.contributions, eager end-of-turn answering, and persona/goal eval simulations.

pipecat-ai 1.9.0

Hume AIModel providerother2026-09-10net-new

Hume publishes a multi-axis voice cloning leaderboard

Public multi-axis cloning scores rather than one Elo. VoxCPM2 led same-speaker (4.21), Cartesia sonic-3.6-beta led naturalness (4.36), Inworld TTS-2 led quality (4.61).

Details1 source · 1 X post

Hume published a Voice Replication Leaderboard measuring how well 11 TTS models clone 25 reference voices across identity, quality, naturalness, and objective similarity.

Voice Replication Leaderboard

New from Hume AI: the Voice Replication Leaderboard. Live now on @huggingface, links

https://x.com/hume_ai/status/2098158670131659137
PolyAIPlatformlaunch2026-09-09

PolyAI launches Wren, an agent that improves dialog agents from live conversations

Moves PolyAI from shipping Dialog-RSN-1 models into continuous agent optimisation on live voice and digital traffic, with named customers (Golden Nugget, Hawksmoor, Simplyhealth) already using it.

Details1 source

PolyAI launched Wren inside Dialog Studio: an agent that reviews every production conversation, proposes fixes and improvements, tests them, and shows metric movement for human approval.

Wren

DeepgramPlatformAPI2026-09-08

Deepgram Voice Agent can defer function calls until end of turn

Closes a speculative-reply footgun where end_call or booking tools could fire before Flux confirmed the user had finished speaking.

Also Think

Details3 sources

Deepgram added defer_until_eot on Voice Agent functions so irreversible tool calls wait for turn confirmation, and emits FunctionCallCancelled when barge-in or a resumed turn voids a speculative call.

Voice Agent defer_until_eot

Connecttelephony, channels1

Daily / PipecatPlatformAPI2026-09-11

Pipecat 1.10.0 fixes WebSocket transport pacing with stream resamplers

Telephony hangup-on-EndFrame bots stop cutting off the last sentence after resample mismatches.

Details3 sources

pipecat-ai 1.10.0 fixes WebSocket transports that treated resampled-empty serializer payloads as unsent, which drained output early and advanced EndFrame too soon.

pipecat-ai 1.10.0

Useagents in market5

ZendeskPlatformlaunch2026-09-10net-new

Zendesk ships native Real-Time Voice Translation for Contact Center

CCaaS embeds bidirectional live translation in the native agent desktop — language routing becomes optional rather than a staffing wall, still EAP-gated.

Also ConnectThink

Details3 sources

Zendesk introduced AI-powered two-way Real-Time Voice Translation built into Zendesk Contact Center so agents and customers can each speak their preferred language on live calls without interpreters.

Real-Time Voice Translation

IDFC FIRST BankEnterprisepartnership2026-09-09net-new

IDFC FIRST Bank and Sarvam open a joint R&D lab for a “self-improving bank”

Indian bank × sovereign AI lab pairing is a concrete enterprise path for continuous post-training under banking controls, not a one-shot model deploy.

Details2 sources · 1 X post

On 9 Sep, IDFC FIRST Bank and Sarvam AI announced a joint R&D lab aimed at banking AI that learns from operational outcomes, with a post-training factory and regulator-grade safety work.

IDFC FIRST Bank × Sarvam AI R&D Lab · vendor Sarvam AI · delivered by Sarvam AI

Parakeet HealthPlatformpartnership2026-09-08net-new

Parakeet Health partners with Qualderm on dermatology patient access

National specialty MSO patient-access wire for a Tarush-listed front-office vendor — confirm channel mix (voice vs digital) on later primary posts.

Details3 sources

Parakeet Health and Qualderm announced a collaboration to expand patient access across Qualderm’s ~160-practice dermatology network by filling capacity after cancellations, missed visits, and overdue follow-ups with conversational agents.

Patient access × Qualderm

KisshtEnterpriselaunch2026-09-08net-new

Kissht’s Ring app: Arrowhead voice agents handle 40% of inbound support

A named lender at 40% inbound containment after two months. Arrowhead claims ~88% resolution on those intents; Kissht did not blog it.

Details2 sources

Indian digital lender Kissht was reported to have Arrowhead voice agents handling 40% of inbound customer-support calls on its Ring app, two months after deployment.

Ring app inbound support (Arrowhead voice agents) · vendor Arrowhead

AOK PLUSEnterpriselaunch2026-09-08net-new

AOK PLUS goes live on NiCE Cognigy and CXone

Regulated EU insurer on NiCE’s sovereign cloud. Vendor-stated 5M+ interactions/year, call acceptance above 95%. Announced by NiCE, not AOK PLUS.

Details1 source

German public health insurer AOK PLUS went live on NiCE Cognigy and CXone, combining AI voice self-service with member-service operations for more than 5 million annual interactions.

Member voice self-service (NiCE Cognigy + CXone) · vendor NiCE