hotwirednews Install bot

Week 16

13–19 Apr 20262026-W16 · 15 filings · 10 companies · 5 net-new

The Orchestrate layer led the week with 4 of 15 filings, from Lua, LiveKit, Cloudflare and Daily / Pipecat.

Hearspeech-to-text1

SensoryModel providermodel release2026-04-15net-new

Sensory ships ultra-compact on-device STT for TFLM and NPUs

Edge STT at 2.7–13 MB with 37 languages targets offline/privacy devices — MAC and SRAM figures are vendor lab specs.

Details1 source

Sensory launched an on-device speech-to-text engine optimised for TensorFlow Lite Micro and NPUs including Arm Ethos-U55, with 37-language support and two footprints: a 2.7 MB domain-specific model and a 13 MB general-purpose model.

On-device STT engine

Speaktext-to-speech2

GoogleModel providermodel release2026-04-15

Google ships Gemini 3.1 Flash TTS with natural-language audio tags

Preview TTS with inline audio tags across 70+ languages and SynthID on every clip — Elo 1,211 is Artificial Analysis via Google's post, treat as third-party but vendor-framed.

Details2 sources

Google introduced Gemini 3.1 Flash TTS in preview on the Gemini API, Google AI Studio, and Vertex AI, adding natural-language audio tags for style, pace, and delivery across 70+ languages with SynthID watermarking on all audio.

Gemini 3.1 Flash TTS

Daily / PipecatPlatformlaunch2026-04-14

Pipecat 1.0.0 adds Mistral Voxtral TTS and NVIDIA cross-sentence stitching

Voxtral TTS and smoother NVIDIA Magpie turns expand open and hyperscaler speak options at GA.

Details3 sources

pipecat-ai 1.0.0 adds MistralTTSService for Voxtral TTS and enhances NvidiaTTSService with cross-sentence stitching within an LLM turn.

pipecat-ai 1.0.0

Showavatars, video3

LemonSliceModel providermodel release2026-04-17

LemonSlice 2.1 Flash targets sub-second avatar time-to-first-byte

Dated LiveKit Agents latency write-up for the pre-CWM stack. The later CWM-1 page also cites 471 ms, but that figure is interrupt frame time, not this Flash TTFT protocol.

Details1 source · 1 X post

LemonSlice 2.1 Flash is a video diffusion-transformer avatar model and inference stack that LemonSlice reports at 471 ms average time-to-first-byte, and 2.04 s average end-to-end when paired with third-party STT, LLM, and TTS.

LemonSlice 2.1 Flash

We did it! We built the fastest interactive avatar model Introducing LemonSlice-2.1 𝘍𝘭𝘢𝘴𝘩 ⚡ Here’s how we did it using @modal and @livekit 👇 (note: it was not easy)

https://x.com/LemonSliceAI/status/2045203651635843378
LiveKitPlatformlaunch2026-04-16

LiveKit Agents 1.5.3 adds Runway Characters avatar plugin

New avatar vendor for Agents. Class-2: folded quiet patch(es) 1.5.2 into this row.

Details2 sources

livekit-agents 1.5.3 adds a Runway Characters avatar plugin.

livekit-agents 1.5.3

Daily / PipecatPlatformAPI2026-04-14

Pipecat 1.0.0 adds LemonSlice avatar connect events

Presence events prevent talking to an avatar that has not joined yet.

Details3 sources

pipecat-ai 1.0.0 adds on_avatar_connected and on_avatar_disconnected on the LemonSlice transport when the avatar joins or leaves the Daily room.

pipecat-ai 1.0.0

ThinkLLMs, speech-to-speech2

LiveKitPlatformlaunch2026-04-16

LiveKit Agents 1.5.3 adds Cerebras LLM and xAI Grok inference

LLM provider cluster for 1.5.3. Class-2: folded quiet patch(es) 1.5.7, 1.5.14 into this row.

Details3 sources

livekit-agents 1.5.3 adds a Cerebras LLM plugin and xAI Grok LLM support on the inference path.

livekit-agents 1.5.3

Daily / PipecatPlatformlaunch2026-04-14

Pipecat 1.0.0 adds Inworld Realtime and WebSocket OpenAI Responses

Realtime Inworld and persistent Responses sockets become default think-layer paths at 1.0.

Details3 sources

pipecat-ai 1.0.0 adds InworldRealtimeLLMService (cascade STT/LLM/TTS with semantic VAD) and makes WebSocket OpenAIResponsesLLMService the Responses API default.

pipecat-ai 1.0.0

Orchestrateframeworks, evals4

LuaPlatformlaunch2026-04-16net-new

Lua ships Spaces supervisor agents to coordinate specialist fleets

Multi-agent supervisor with voice support on the same Agent OS that later added Browser/WhatsApp/Phone/Meetings transports — orchestration first, channel second.

Also Use

Details1 source

Lua launched Spaces: supervisor agents that reason over a fleet of specialised agents, delegate in one conversation, and synthesise a single reply across web, WhatsApp, Slack, Teams, and voice.

Spaces

LiveKitPlatformlaunch2026-04-16

LiveKit Agents 1.5.3 reuses realtime sessions and adds InstructionParts workflows

Realtime handoff reuse is the orchestration headline.

Details1 source

livekit-agents 1.5.3 reuses realtime sessions across agent handoffs when supported, and adds beta InstructionParts plus ToolSearchToolset and ToolProxyToolset.

livekit-agents 1.5.3

CloudflarePlatformlaunch2026-04-15net-new

Cloudflare adds experimental @cloudflare/voice to Agents SDK

Puts voice on the same Durable Object agent as text and tools without a separate framework — experimental, with Workers AI defaults and a Twilio adapter caveat on mulaw audio.

Also HearSpeakConnect

Details2 sources

Cloudflare released an experimental @cloudflare/voice package for the Agents SDK so the same Durable Object agent can take real-time voice over its existing WebSocket, with built-in Workers AI Flux/Nova 3 STT and Aura TTS plus a Twilio phone adapter.

@cloudflare/voice

Daily / PipecatPlatformlaunch2026-04-14

Pipecat hits 1.0.0 with Settings-pattern services and major removals

1.0 locks the builder-facing service and frame contract the rest of 2026 builds on.

Details3 sources

pipecat-ai 1.0.0 stabilises the Settings-pattern service APIs, async function calls with cancel_on_interruption=False, LLMMessagesTransformFrame, and removes deprecated pre-1.0 surfaces.

pipecat-ai 1.0.0

Useagents in market3

KotobaModel providerpartnership2026-04-17net-new

Kotoba joins SoftBank Daredemo AI and ships Japanese–Korean S2S interpretation

Consumer/carrier distribution for Kotoba plus a named East-Asia S2S language pair beyond the May API alpha and August Agentforce JP filings.

Also ThinkSpeak

Details2 sources

Kotoba integrated its realtime interpretation app into SoftBank's Daredemo AI on 17 April 2026 and officially released a Koto-powered Japanese–Korean speech-to-speech model.

SoftBank Daredemo AI + Japanese–Korean S2S

DeepLPlatformlaunch2026-04-16

DeepL launches Voice-to-Voice real-time spoken translation suite

Speech-to-speech translation suite with staggered GA/early-access SKUs and 40+ languages — Slator 96% preference is DeepL-commissioned.

Also ThinkConnect

Details2 sources

DeepL launched DeepL Voice-to-Voice for live spoken translation across meetings, mobile/web conversations, group QR sessions, and a Voice-to-Voice API, expanding DeepL Voice to over forty languages including all 24 official EU languages.

DeepL Voice-to-Voice

Dairy QueenEnterprisepartnership2026-04-16net-new

Dairy Queen expands Presto Voice AI drive-thru pilots to select franchisees

Named national QSR operator go-live path — corporate test into select franchisee expansion — for Presto Voice rather than an unnamed pilot.

Details1 source

American Dairy Queen Corporation partnered with Presto on drive-thru Voice AI after corporate-store tests, expanding the pilot to selected franchisees across the US.

Drive-thru Voice AI (Presto) · vendor Presto