hotwirednews Install bot

Week 3

12–18 Jan 20262026-W03 · 14 filings · 8 companies · 7 net-new

The Hear, Speak and Orchestrate layers led the week, with 3 filings each.

Hearspeech-to-text3

Daily / PipecatPlatformAPI2026-01-13

Pipecat 0.0.99 lets Gladia and Deepgram drive turn boundaries from STT

Server-side STT endpointing becomes a first-class alternative to pipeline VAD for turn control.

Details3 sources

pipecat-ai 0.0.99 adds enable_vad on GladiaSTTService, should_interrupt on Deepgram and Speechmatics STT, Speechmatics split_sentences, and use_ssl on NVIDIA STT/TTS.

pipecat-ai 0.0.99

DeepgramModel providerfunding2026-01-13net-new

Deepgram raises $130M Series C at a $1.3B valuation and acquires OfOne

Moves Deepgram's first_seen earlier than the April Flux filing and records the OfOne restaurant acquihire on the same primary — a platform raise, not a model ship.

Also Speak

Details2 sources

Deepgram announced $130 million in Series C funding led by AVP at a $1.3 billion valuation, and the same release says it acquired restaurant voice startup OfOne.

Deepgram platform

AmazonPlatformmodel release2026-01-13net-new

Amazon Lex adds neural ASR model for English voice bots

Contact-centre Lex bots get a Neural ASR toggle for accented English on 13 January — still Lex-scoped, not a standalone Transcribe SKU.

Also Connect

Details1 source

Amazon Lex launched a neural automatic speech recognition model for English locales that improves recognition of conversational speech, non-native speakers, and regional accents inside Lex voice bots.

Amazon Lex neural ASR (English)

Speaktext-to-speech3

LiveKitPlatformlaunch2026-01-14net-new

LiveKit Agents 1.3.11 adds Gemini streaming TTS path work

Keeps TTS separate from hear and orchestrate clusters.

Details1 source

livekit-agents 1.3.11 includes Gemini streaming TTS support on the speak path alongside provider websocket model query params.

livekit-agents 1.3.11

Daily / PipecatPlatformAPI2026-01-13

Pipecat 0.0.99 adds Azure TTS word timestamps and Cartesia pronunciation dicts

Word sync and pronunciation control matter for captioning and branded voice quality in cascaded bots.

Details3 sources

pipecat-ai 0.0.99 adds word-level timestamps on AzureTTSService, pronunciation_dict_id for Cartesia TTS, AudioContextTTSService, and Inworld TTS keepalive.

pipecat-ai 0.0.99

KyutaiModel providermodel release2026-01-13net-new

Kyutai open-sources Pocket TTS: 100M-parameter CPU voice cloning

Open MIT TTS that clones from ~5 s of audio and stays real-time on an Intel Ultra 7 / Apple M3 laptop — vendor benches versus F5-TTS and Kokoro, English-public-data training only.

Details3 sources

Kyutai released Pocket TTS, a 100-million-parameter MIT-licensed text-to-speech model with few-second voice cloning that runs faster than real time on laptop CPUs, trained only on public English speech data.

Pocket TTS

Showavatars, video1

Daily / PipecatPlatformlaunch2026-01-13

Pipecat 0.0.99 wires HeyGen LiveAvatar into HeyGenTransport

Avatar bots can use LiveAvatar without a parallel custom video path.

Details3 sources

pipecat-ai 0.0.99 adds HeyGen LiveAvatar API support on HeyGenTransport for realtime avatar video in Daily rooms.

pipecat-ai 0.0.99

ThinkLLMs, speech-to-speech2

NVIDIAModel providermodel release2026-01-15net-new

NVIDIA open-sources PersonaPlex full-duplex speech model with role and voice prompts

Open full-duplex S2S finally takes text role prompts and voice clones together — FullDuplexBench wins are NVIDIA's own charts.

Also Speak

Details3 sources

NVIDIA ADLR released PersonaPlex, a 7B Moshi-based full-duplex speech-to-speech model that takes a voice prompt plus a text role prompt so agents keep a chosen persona while listening and speaking concurrently.

PersonaPlex

Daily / PipecatPlatformlaunch2026-01-13

Pipecat 0.0.99 adds Grok Realtime speech-to-speech

Grok Voice joins the Pipecat realtime service set beside OpenAI and Gemini Live paths.

Details3 sources

pipecat-ai 0.0.99 adds GrokRealtimeLLMService for xAI's Grok Voice Agent API with realtime audio streaming and function calling.

pipecat-ai 0.0.99

Orchestrateframeworks, evals3

LiveKitPlatformlaunch2026-01-14net-new

LiveKit Agents 1.3.11 adds EndCallTool and MCP HTTP transport options

First 2026 Python Agents tag; MCP and EndCallTool are the orchestration headline.

Details1 source

livekit-agents 1.3.11 ships EndCallTool, standardises the Tool interface, and adds allowed_tools and transport_type on MCPServerHTTP.

livekit-agents 1.3.11

TavusModel providermodel release2026-01-13net-new

Tavus ships Sparrow-1 conversational-flow model for realtime turn-taking

Modular ASR→LLM→TTS stacks get a dedicated timing layer — Tavus's 55ms p50 / zero-interruption bench is on a 28-sample internal set.

Also Think

Details1 source

Tavus released Sparrow-1, an audio-native conversational-flow model that predicts floor ownership so agents know when to listen, wait, or speak, instead of relying on silence-based endpoint detection.

Sparrow-1

Daily / PipecatPlatformlaunch2026-01-13

Pipecat 0.0.99 introduces composable user turn and mute strategies

Turn-taking becomes an ordered strategy list — the foundation later 2026 releases extend with incomplete-turn, eager EOT, and wake-phrase behaviour.

Details3 sources

pipecat-ai 0.0.99 adds ordered user turn start/stop strategies, user mute strategies, UserTurnController, and UserTurnProcessor, replacing transport-level turn knobs.

pipecat-ai 0.0.99

Connecttelephony, channels1

Daily / PipecatPlatformAPI2026-01-13

Pipecat 0.0.99 adds Vonage Audio Connector serializer

Expands Pipecat's telephony connector surface beyond Daily, Twilio, and Telnyx serializers.

Details3 sources

pipecat-ai 0.0.99 adds VonageFrameSerializer for the Vonage Video API Audio Connector WebSocket protocol.

pipecat-ai 0.0.99

Useagents in market1

SpeechifyPlatformlaunch2026-01-12net-new

Speechify launches Voice AI Assistant on iOS

Consumer voice assistant lands on iOS after web/Chrome — phone-call automation is still roadmap, not this cut.

Also Think

Details1 source

Speechify released its Voice AI Assistant on iOS, extending the Chrome and web assistant so users can ask questions, browse, run multi-turn voice dialogue, and trigger actions such as creating a podcast or opening a book by voice.

Speechify Voice AI Assistant (iOS)