hotwirednews Install bot

Week 20

11–17 May 20262026-W20 · 9 filings · 6 companies · 3 net-new

The Speak, Orchestrate and Use layers led the week, with 2 filings each.

Hearspeech-to-text1

Daily / PipecatPlatformAPI2026-05-14

Pipecat 1.2.0 expands ElevenLabs STT keyterm steering

Domain vocabulary steering reduces proper-noun errors on Scribe pipelines.

Details3 sources

pipecat-ai 1.2.0 improves ElevenLabs STT services so Scribe callers can bias transcription with keyterms.

pipecat-ai 1.2.0

Speaktext-to-speech2

Daily / PipecatPlatformAPI2026-05-14

Pipecat 1.2.0 adds Cartesia buffer delay and Deepgram MIP opt-out

Latency versus prosody trade-offs become per-bot settings rather than provider defaults only.

Details3 sources

pipecat-ai 1.2.0 adds max_buffer_delay_ms on CartesiaTTSService and mip_opt_out on Deepgram TTS services.

pipecat-ai 1.2.0

Resemble AIModel providermodel release2026-05-13net-new

Resemble ships DramaBox prompt-directed expressive TTS

Screenplay-style direction plus default watermarking targets game/audiobook/agent delivery that flat TTS still misses — English-only by design, not a multilingual race.

Details1 source

Resemble released DramaBox, an English-only 3.3B diffusion-transformer TTS directed by plain-language screenplay prompts, with optional 10-second voice cloning, ~2.5 s generation on a warm H100, and Resemble Watermarker embedded on every output.

DramaBox

Showavatars, video1

Boson AIModel providermodel release2026-05-13

Boson AI launches Higgs Avatar for real-time talking faces

Adds a face layer to Boson's own STT/TTS stack for production voice agents — private preview only, eight sessions per H100 is the cost claim to watch.

Also Speak

Details1 source

Boson AI introduced Higgs Avatar, a real-time avatar foundation model that generates an expressive speaking face from a single still image, co-designed with Higgs TTS, streaming frames at ~16 ms on one H100 hosting up to eight concurrent conversations.

Higgs Avatar

Orchestrateframeworks, evals2

Daily / PipecatPlatformlaunch2026-05-14

Pipecat 1.2.0 adds incomplete-turn and deferred stop strategies

Incomplete-turn handling moves from an LLM mixin into composable turn strategies. Class-2: folded quiet patch(es) 1.2.1 into this row.

Details5 sources

pipecat-ai 1.2.0 adds FilterIncompleteUserTurnStrategies, DeferredUserTurnStopStrategy, ExternalUserTurnCompletionStopStrategy, and LLMTurnCompletionUserTurnStopStrategy.

pipecat-ai 1.2.0

LiveKitPlatformlaunch2026-05-13

LiveKit Agents 1.5.9 introduces Answering Machine Detection

Outbound call classification is a new orchestration capability.

Details1 source

livekit-agents 1.5.9 introduces Answering Machine Detection so outbound calls can classify person, voicemail, IVR, or unreachable numbers.

livekit-agents 1.5.9

Connecttelephony, channels1

Daily / PipecatPlatformAPI2026-05-14

Pipecat 1.2.0 adds RTVI UI Agent Protocol messages

Voice bots can observe and cancel UI jobs over the same RTVI channel as audio.

Details3 sources

pipecat-ai 1.2.0 adds first-class RTVI ui-event, ui-snapshot, and ui-cancel-task client-to-server messages for UI-driven agents.

pipecat-ai 1.2.0

Useagents in market2

KolsetuPlatformlaunch2026-05-13net-new

Kolsetu opens Elba Self-Serve voice AI to EU builders

EU self-serve voice orchestration with compliance as the opener — lowers the bar for agencies and SMBs that refuse US-only stacks, while keeping an enterprise services lane.

Also OrchestrateConnect

Details1 source

Kolsetu opened Elba Self-Serve on 13 May for European builders: GDPR-oriented voice-agent platform with provider-flexible models, MCP server connections, no credit card to start, and an enterprise path for mission-critical deployments.

Elba Self-Serve

Yellow.aiPlatformlaunch2026-05-11net-new

Yellow.ai launches Nexus Vox zero-hop enterprise voice

Enterprise CX platform ships voice as a native runtime layer rather than another bolted STT/TTS pair — 500+ language claim and sub-400 ms e2e are vendor figures to pressure-test.

Also ThinkSpeakConnect

Details1 source

Yellow.ai launched Nexus Vox inside its Nexus agentic platform: a single-runtime voice stack claiming sub-400 ms end-to-end latency, 500+ languages/dialects, 10-second voice cloning, real-time sentiment, and direct CRM/ticketing orchestration under one SLA.

Nexus Vox