LiveKit Agents 1.5.15 adds Cartesia ink-2 STT
Cartesia STT model addition. Class-2: folded quiet patch(es) 1.5.17 into this row.
Details2 sources
livekit-agents 1.5.15 adds Cartesia ink-2 STT.
livekit-agents 1.5.15
The Hear, Speak, Think and Orchestrate layers led the week, with 2 filings each.
Cartesia STT model addition. Class-2: folded quiet patch(es) 1.5.17 into this row.
livekit-agents 1.5.15 adds Cartesia ink-2 STT.
livekit-agents 1.5.15
Cartesia's turn-aware ASR becomes selectable beside Deepgram Flux and AssemblyAI U3 Pro.
pipecat-ai 1.3.0 adds CartesiaTurnsSTTService for Cartesia Streaming ASR with turn signals, plus max_endpoint_delay_ms on SonioxSTTService.
pipecat-ai 1.3.0
Keeps Rime users on current model names after Arcana defaults in earlier releases.
pipecat-ai 1.3.0 adds the Rime coda model to RimeTTSService and RimeHttpTTSService.
pipecat-ai 1.3.0
Performance-conditioned multilingual speech — not flat TTS from a transcript — is the gap studios and creators still pay humans to close; API still sales-gated at launch.
ElevenLabs launched Dubbing v2, which conditions on the original speaker's performance to carry tone, pacing, and emotion across 90+ languages with sync-aware translation, available in ElevenCreative and ElevenProductions with API access still forthcoming.
Dubbing v2
Diffusion LLM option for builders comparing cascade latency versus autoregressive peers.
pipecat-ai 1.3.0 adds InceptionLLMService for Inception's Mercury 2 diffusion reasoning model.
pipecat-ai 1.3.0
First developer-facing Kotoba surface in the corpus; pairs with the June seed and August Agentforce Japan filing.
Also HearSpeak
Kotoba Technologies released an alpha API and Python SDK exposing speech-to-speech translation plus streaming STT and TTS, aimed at builders evaluating East Asian voice workloads.
Kotoba API & SDK Alpha
Subagents and UI workers become a core runtime pattern rather than parallel pipelines glued by hand.
pipecat-ai 1.3.0 makes pipelines multi-agent compatible by default and adds the pipecat.workers framework including UIWorker for observing and driving client web UIs.
pipecat-ai 1.3.0
Splits the stack so ElevenLabs owns the speech loop and builders keep their own LLM — the missing middle between raw TTS/STT APIs and a fully hosted ElevenAgents runtime.
Also HearSpeak
ElevenLabs released Speech Engine, a WebSocket path that handles STT, turn-taking, TTS, and browser playback while the developer's server owns the LLM logic and streams response text — an alternative to fully hosted ElevenAgents.
Speech Engine
Vonage WebRTC joins Daily, LiveKit, and SmallWebRTC as a first-party transport.
pipecat-ai 1.3.0 adds VonageVideoConnectorTransport for realtime Vonage WebRTC sessions.
pipecat-ai 1.3.0