hotwirednews Install bot

Week 18

27 Apr – 3 May 20262026-W18 · 9 filings · 6 companies · 2 net-new

The Think layer led the week with 3 of 9 filings, from Sakana AI, AssemblyAI and Microsoft.

Hearspeech-to-text2

DeepgramModel providermodel release2026-04-29

Deepgram GA Flux STT Multilingual across 10 languages

Single flux-general-multi model replaces per-language STT stacks with mid-call code-switching — WER and end-of-turn latency benches are Deepgram's own.

Details3 sources

Deepgram generally available Flux STT Multilingual (flux-general-multi), a conversational speech recognition model covering English, Spanish, French, German, Hindi, Russian, Portuguese, Japanese, Italian, and Dutch with native code-switching, turn detection, and interruption handling on one streaming connection.

Flux STT Multilingual

Daily / PipecatPlatformlaunch2026-04-27

Pipecat 1.1.0 adds xAI and Mistral realtime STT

xAI and Mistral join the streaming STT catalogue beside Deepgram Flux multilingual hints.

Details3 sources

pipecat-ai 1.1.0 adds XAISTTService and MistralSTTService (Voxtral Realtime), plus Deepgram Flux language_hints and Sarvam saaras:v3 server-side VAD tuning.

pipecat-ai 1.1.0

Speaktext-to-speech1

Daily / PipecatPlatformlaunch2026-04-27

Pipecat 1.1.0 adds xAI and Soniox streaming TTS

Pairs with XAISTTService so builders can run an all-xAI cascade.

Details3 sources

pipecat-ai 1.1.0 adds XAITTSService for xAI WebSocket TTS and SonioxTTSService for persistent WebSocket speech synthesis.

pipecat-ai 1.1.0

ThinkLLMs, speech-to-speech3

Sakana AIModel providerpaper2026-05-03net-new

Sakana AI introduces KAME tandem speech-to-speech with live LLM oracles

Tandem Moshi-front plus LLM-oracle design targets cascade median latency near 2.1 s without a shipped product API — research demo only as of the 3 May coverage.

Details2 sources

Sakana AI published KAME (Knowledge-Access Model Extension), a tandem architecture that keeps a Moshi-style front-end speech-to-speech path speaking immediately while a back-end LLM streams improving text oracles into the conversation in real time.

KAME

AssemblyAIPlatformAPI2026-04-29net-new

AssemblyAI launches end-to-end Voice Agent API on one WebSocket

Flat $4.50/hr all-in WebSocket competes with multi-vendor STT/LLM/TTS metering — STT accuracy claims and the 76% survey figure are vendor-published.

Also HearSpeak

Details1 source

AssemblyAI released a Voice Agent API that runs speech understanding, LLM reasoning, and voice generation behind a single WebSocket at a flat $4.50 per hour, with server-side turn detection, tool calling, live config updates, and session resumption.

Voice Agent API

MicrosoftPlatformlaunch2026-04-27

Microsoft GA real-time voice agents in Copilot Studio for Dynamics 365

Premium S2S mode inside Dynamics 365 Contact Center is GA in North America first — Teams Phone and broader Copilot Studio channels stay roadmap, and the Fortune 500 agent claim is vendor-stated.

Also UseConnect

Details1 source

Microsoft made real-time voice agents generally available in Copilot Studio for Dynamics 365 Contact Center in North America, adding a premium interruptible speech-to-speech mode with live reasoning on top of existing voice automation.

Copilot Studio real-time voice agents

Orchestrateframeworks, evals1

Daily / PipecatPlatformAPI2026-04-27

Pipecat 1.1.0 adds tool_resources for shared app state in tool handlers

Removes a common pattern of smuggling clients through closures or module globals.

Details3 sources

pipecat-ai 1.1.0 adds tool_resources on PipelineTask and FunctionCallParams so handlers receive DB handles, clients, and other app state without globals.

pipecat-ai 1.1.0

Connecttelephony, channels1

Daily / PipecatPlatformAPI2026-04-27

Pipecat 1.1.0 adds Daily DTMF send and screenVideo output

Outbound DTMF and screen-share video destinations cover IVR navigation and screen-share agent UIs.

Details3 sources

pipecat-ai 1.1.0 adds DailyTransport.send_dtmf(), DailyOutputDTMF frames, multi-key OutputDTMFFrame.buttons, and Daily screenVideo destination support.

pipecat-ai 1.1.0

Useagents in market1

LuaPlatformlaunch2026-04-30

Lua ships Voice across Browser, WhatsApp, Phone, and Meetings

Same agent persona across four voice transports including meeting participation roles — Meetings is early-access, and the 5,000-agent figure is vendor-claimed for the prior text base.

Also ConnectThink

Details1 source

Lua launched Lua Voice so one agent configuration runs real-time interruptible voice on Browser, WhatsApp, Phone, and Meetings (Google Meet, Zoom, Microsoft Teams), with the same persona, tools, and governance as its text agents.

Lua Voice