Daily / Pipecat@pipecat_ai
Pipecat 1.10.0 adds Meta STT and Speechmatics Agent STT
pipecat-ai 1.10.0 adds MetaSTTService for muse-voice-transcribe-1.0, switches SpeechmaticsSTTService onto Speechmatics Agent STT at /v2/agent with linden-1 as the default model, and removes LmntTTSService after LMNT shut down.
Same week as LiveKit Agents 1.8.1 DuplexModel and OpenAI's 10 Sep GPT-Live-1 API. Pipecat's path is a pipeline service, not LiveKit's DuplexModel class. Filed against the CHANGELOG, not a blog post.
Tavus@tavus
Tavus Phoenix-4.5 renders full upper-body PALs at 134 ms audio-to-video
Tavus shipped Phoenix-4.5 on 10 Sep: a full-frame generative renderer that moves head, shoulders, posture, and torso with speech, at 134 ms audio-to-video.
Realtime avatar stacks that only animate a face leave body motion static; whole-frame generation plus sub-150 ms A2V is the bar conversational video agents now have to clear.
OpenAI@OpenAI
GPT-Live-1 in the API: full-duplex voice at $0.05/min
OpenAI brought ChatGPT’s full-duplex GPT-Live-1 voice model to the developer API at $0.05/min for the voice layer, with backend models billed separately.
Community cite: ~0.798s turn-taking versus 1.41s for GPT-Realtime-2.1. Backend model billed separately.
LiveKit@livekit
LiveKit Agents 1.8.1 adds DuplexModel for GPT-Live
livekit-agents 1.8.1 introduces DuplexModel so a session can run a full-duplex speech model, with GPTLiveModel as the first implementation.
SDK counterpart of OpenAI's 10 Sep GPT-Live-1 API. Expressive mode is already filed from the 12 Aug blog; that landed in 1.7.0 on 20 Aug. Do not re-file 1.7.0.
GPT-Live-1 @openaidevs is now supported in LiveKit Agents. It's full-duplex, so it listens and speaks at the same time instead of taking turns, and it hands reasoning and tool calls off to a backend model so a slow lookup doesn't leave your agent sitting in dead air.
https://x.com/livekit/status/2098126102052905001
Hume AI@hume_ai
Hume publishes a multi-axis voice cloning leaderboard
Hume published a Voice Replication Leaderboard measuring how well 11 TTS models clone 25 reference voices across identity, quality, naturalness, and objective similarity.
Public multi-axis cloning scores rather than one Elo. VoxCPM2 led same-speaker (4.21), Cartesia sonic-3.6-beta led naturalness (4.36), Inworld TTS-2 led quality (4.61).
New from Hume AI: the Voice Replication Leaderboard. Live now on @huggingface, links
https://x.com/hume_ai/status/2098158670131659137
PolyAI@polyaivoice
PolyAI launches Wren, an agent that improves dialog agents from live conversations
PolyAI launched Wren inside Dialog Studio: an agent that reviews every production conversation, proposes fixes and improvements, tests them, and shows metric movement for human approval.
Moves PolyAI from shipping Dialog-RSN-1 models into continuous agent optimisation on live voice and digital traffic, with named customers (Golden Nugget, Hawksmoor, Simplyhealth) already using it.
IDFC FIRST Bank@IDFCFIRSTBank
IDFC FIRST Bank and Sarvam open a joint R&D lab for a “self-improving bank”
On 9 Sep, IDFC FIRST Bank and Sarvam AI announced a joint R&D lab aimed at banking AI that learns from operational outcomes, with a post-training factory and regulator-grade safety work.
Indian bank × sovereign AI lab pairing is a concrete enterprise path for continuous post-training under banking controls, not a one-shot model deploy.
HeyGen@HeyGen
HeyGen Professional Voice Clone trains a dedicated voice from 20+ minutes of audio via API
HeyGen opened Professional Voice Clone on the API on 9 Sep: a paid per-slot HeyGen Voice adapter trained on 1–10 recordings totaling at least 20 minutes of one speaker.
Avatar video quality has outrun voice fidelity; a studio-grade clone behind the same API as LiveAvatar closes that gap for builders who already have 20 minutes of clean speaker audio.
Gradium@GradiumAI
Gradium retrains default TTS after Coval starts scoring leading silence in TTFA
After Coval’s June perceived-TTFA change exposed ~225 ms median leading silence in Gradium’s stream, Gradium shipped a new default that emits audible audio in the first 10 ms frame (median TTFA ~214 ms).
Labs that game time-to-first-byte with silent frames get caught when evals score what a caller hears; Gradium’s response shows the metric now steers shipping.
Kissht@Kissht_EMI
Kissht’s Ring app: Arrowhead voice agents handle 40% of inbound support
Indian digital lender Kissht was reported to have Arrowhead voice agents handling 40% of inbound customer-support calls on its Ring app, two months after deployment.
A named lender at 40% inbound containment after two months. Arrowhead claims ~88% resolution on those intents; Kissht did not blog it.
AOK PLUS
AOK PLUS goes live on NiCE Cognigy and CXone
German public health insurer AOK PLUS went live on NiCE Cognigy and CXone, combining AI voice self-service with member-service operations for more than 5 million annual interactions.
Regulated EU insurer on NiCE’s sovereign cloud. Vendor-stated 5M+ interactions/year, call acceptance above 95%. Announced by NiCE, not AOK PLUS.