hotwirednews Install bot

Week 35

24–30 Aug 20262026-W35 · 22 filings · 14 companies · 5 net-new

The Orchestrate layer led the week with 7 of 22 filings, from Sesame, Tavus, Cartesia, Daily / Pipecat and 2 more.

Hearspeech-to-text6

FermionModel providermodel release2026-08-28net-new

Fermion ships Phonon-1 open-weight English ASR in a 415 MB download

Puts a 415 MB Apache-2.0 English ASR path claiming better accuracy than a Whisper twice its size on laptop runtimes — vendor LibriSpeech and MacBook Air figures, not an independent audit — alongside same-day Micro/Big Hub siblings.

Details3 sources · 1 X post

Fermion released Phonon-1, a 782 million-parameter English speech-recognition model that downloads in 415 MB, claims higher accuracy than a Whisper checkpoint twice its size, and transcribes about an hour of audio in roughly two minutes on a MacBook Air under Apache 2.0.

Phonon-1

Introducing Phonon-1: a new frontier in speech recognition per byte. A 782 million param model in a 415 MB download, more accurate than a Whisper 2x its size. Transcribe an hour of audio in about two minutes on a MacBook Air. Open weights, Apache 2.0, today. Tags @fermion_ai.

https://x.com/yoitsmanan/status/2093388796272562222
DeepgramModel providerAPI2026-08-28

Deepgram adds ForceEndTurn and flexible Flux turn-taking

Push-to-talk, DTMF, and external VAD stacks can own turn boundaries on Flux without reconnecting — a control plane most STT APIs still lack.

Also Orchestrate

Details3 sources

Deepgram gave Flux STT and Voice Agent clients ForceEndTurn plus mid-stream eot_threshold control so builders can run automatic, semi-manual, or fully manual turn endings, including Bring Your Own Turn Detection.

Flux ForceEndTurn

LiveKitPlatformlaunch2026-08-27

LiveKit Agents 1.7.1 adds Sarvam realtime STT, Palabra STT, and Gemini transcribe-live

Large STT provider wave including Sarvam v4.

Details1 source

livekit-agents 1.7.1 adds Sarvam realtime streaming STT with saaras:v4 default, Palabra STT, Gemini 3.5 transcribe-live, and Speechmatics Linden on inference.

livekit-agents 1.7.1

Daily / PipecatPlatformAPI2026-08-26

Pipecat 1.8.0 adds Google STT v2 adaptation

Phrase-set adaptation is the usual lever for proper nouns on Google Chirp pipelines.

Details3 sources

pipecat-ai 1.8.0 adds Speech-to-Text v2 adaptation support on GoogleSTTService so recognition can be biased with phrase sets and models.

pipecat-ai 1.8.0

GoogleModel providermodel release2026-08-26

Gemini 3.5 Transcribe replaces Chirp 3 for many developer workflows

Google’s best ASR on the same Live API path as Gemini voice. Vendor: ~70% faster time-to-final versus Chirp 3. Artificial Analysis WER ~4.0% streaming / 2.6% non-streaming.

Details2 sources · 1 X post

Google launched Gemini 3.5 Transcribe for smart STT in live streaming and pre-recorded modes, replacing Chirp 3 for many developer workflows.

Gemini 3.5 Transcribe

We're introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet

https://x.com/Google/status/2092659278632894576
SukiPlatformlaunch2026-08-25net-new

Suki ships standalone AI Dictation inside Epic and MEDITECH

Unbundles AI dictation from ambient ACI — modular buy for Epic/MEDITECH shops and a partner-embeddable speech path.

Also Use

Details2 sources

Suki launched Suki Dictation, an AI verbatim dictation product that health systems can buy alone or with Suki ambient documentation, natively inside Epic and MEDITECH, with APIs/SDKs for embedded partners.

Suki Dictation

Speaktext-to-speech3

LiveKitPlatformlaunch2026-08-27

LiveKit Agents 1.7.1 adds ElevenLabs v3 dialogue TTS, Speechify timestamps, and Palabra TTS

ElevenLabs v3 dialogue and Speechify timestamps.

Details1 source

livekit-agents 1.7.1 streams eleven_v3/eleven_v3_conversational via text-to-dialogue, adds Speechify timestamp streaming, and Palabra TTS.

livekit-agents 1.7.1

CartesiaModel providermodel release2026-08-27

Cartesia Sonic-3.6: streaming TTS, 44 languages, sub-90ms TTFA

44-language streaming TTS with a sub-90ms TTFA claim. Treat the Artificial Analysis #1 as vendor-cited. Adds Odia, Urdu, and Hinglish code-switch.

Details2 sources · 1 X post

Cartesia shipped Sonic-3.6, a streaming TTS model claiming #1 on both Artificial Analysis speech arenas across 44 languages with sub-90ms time-to-first-audio.

Sonic-3.6

Sonic-3.6 is now generally available. In January we made a bet: stop tuning the existing paradigm, rebuild from the architecture up.

https://x.com/cartesia/status/2093017207185821705
Smallest AIModel providermodel release2026-08-26net-new

Smallest AI ships Lightning V3.1 conversational TTS

Gives the Waves stack a dated V3.1 TTS drop builders can pin, after Smallest sat on the roster with an empty announcement list.

Details2 sources

Smallest AI launched Lightning V3.1, a 44.1 kHz conversational text-to-speech model with voice cloning from 5–15 seconds of audio across 15 languages and mid-sentence language switching.

Lightning V3.1

ThinkLLMs, speech-to-speech3

LiveKitPlatformlaunch2026-08-27

LiveKit Agents 1.7.1 adds Sarvam LLM models including sarvam-105b-conversations

Sarvam and xAI LLM catalogue.

Details1 source

livekit-agents 1.7.1 adds Sarvam LLM support for glm5.2, gemma4, and sarvam-105b-conversations, plus xAI inference LLM models.

livekit-agents 1.7.1

Druid AIPlatformlaunch2026-08-27

Druid Voice Speech-to-Speech technology preview with sub-second duplex

Moves Druid from pipelined Voice GA (W29) toward native S2S/duplex — think-layer preview alongside Gemini Live BYOS option.

Also SpeakHear

Details1 source

Druid 9.30 introduced a technology preview of Druid Voice Speech-to-Speech with sub-second latency, barge-in, mid-sentence language switching across 78 languages, and choice of Druid or Google Gemini Live speech services.

Druid Voice Speech-to-Speech · vendor Google

Daily / PipecatPlatformlaunch2026-08-27

Pipecat releases PhoneLLM Alpha 1 and PhoneBench

Small open MoE for phone-agent tool calls, with a published cost/latency bench. Example self-host cost ~$0.0025/min at high concurrency on B200 via Modal.

Also Orchestrate

Details1 source

Pipecat released an open-weights (BSD) voice-agent LLM fine-tune of NVIDIA Nemotron 3 Nano, plus PhoneBench, a phone-agent eval for accuracy, style, latency, and cost.

PhoneLLM Alpha 1

Orchestrateframeworks, evals7

SesamePlatformlaunch2026-08-28net-new

Sesame open-sources TurnBench for spoken turn-taking

Gives builders a shared EOT/interruption leaderboard beyond Switchboard heuristics — the same board Tavus Sparrow-2 later claims to lead on the public dev split.

Details4 sources

Sesame released TurnBench, an open benchmark for end-of-turn and interruption detection on 30 hours of triple-annotated dual-channel conversations, with a public leaderboard, viewer, and self-serve scoring.

TurnBench

TavusModel providermodel release2026-08-27

Tavus ships Sparrow-2 whole-scene conversational understanding

Moves Tavus turn-taking from Sparrow-1 endpoint detection to whole-room conversational state — noisy retail/café deployments become the product claim, with TurnBench-leading but still preliminary scores.

Also Think

Details4 sources

Tavus released Sparrow-2, an audio-native streaming conversational-understanding model that jointly models turn-taking, interruptions, backchannels, and the acoustic scene at a 10 ms frame rate, now GA on PALs and APIs.

Sparrow-2

CartesiaPlatformlaunch2026-08-27

Cartesia launches Managed Agents on Ink and Sonic

Puts Cartesia's own STT/TTS stack under a hosted agent runtime — a lighter path than self-hosted Line for teams that only need Ink+Sonic plus tools.

Also SpeakHear

Details3 sources

Cartesia shipped Managed Agents so teams can run a production voice agent from instructions, a voice, an LLM, and tools without hosting their own backend, with Cartesia running the call on Ink STT and Sonic TTS.

Cartesia Managed Agents

Daily / PipecatPlatformlaunch2026-08-26

Pipecat 1.8.0 adds Context Hub, MCP helpers, and proposed turn frames

Proposed turns, MCP argument injection, and Context Hub tighten the builder loop without a new STT/TTS vendor.

Details3 sources

pipecat-ai 1.8.0 integrates Context Hub for coding-agent docs, adds MCPClient(tools_arguments=...), ProposedUserStarted/StoppedSpeakingFrame, and ErrorCategory on failures.

pipecat-ai 1.8.0

NiCE CognigyPlatformlaunch2026-08-26net-new

NiCE Cognigy Agentic Building: coding agents ship CX voice agents over MCP

CCaaS is exposing the agent lifecycle to coding agents in the IDE, with the visual UI kept as the inspection surface.

Details1 source

NiCE Cognigy launched Agentic Building via an open-source Agent Plugin so coding assistants (Claude Code, Codex, Gemini CLI, Cursor) can create, test, and deploy Cognigy voice and chat agents from the IDE.

Agentic Building / Agent Plugin

ElevenLabsPlatformlaunch2026-08-24

ElevenLabs makes ElevenAgents Procedures generally available

Lets one agent switch task packs mid-call without a prompt rewrite — useful when the same number handles billing, scheduling, and escalations.

Details2 sources

ElevenLabs opened Procedures on ElevenAgents: task-specific free-form or structured instruction sets the agent loads mid-call instead of stuffing every branch into the system prompt.

ElevenAgents Procedures

ElevenLabsPlatformlaunch2026-08-24

ElevenLabs CLI v1: pull/push agents as code

Agents as versioned files with dry-run and schema checks. Aimed at fleets, not a one-off agent in the web UI.

Details1 source · 1 X post

ElevenLabs released an agents-first CLI covering the full API plus agents-as-code pull/push workflows with --schema and --dry-run.

ElevenLabs CLI v1

Today we are releasing ElevenLabs CLI v1, which brings the entire ElevenLabs API into the terminal.

https://x.com/ElevenLabs/status/2091916523132616961

Connecttelephony, channels1

Daily / PipecatPlatformAPI2026-08-26

Pipecat 1.8.0 adds LiveKit to the development runner

Local LiveKit bots no longer need a custom runner fork.

Details3 sources

pipecat-ai 1.8.0 adds LiveKit as a -t livekit transport option in the development runner and matching POST transport selection.

pipecat-ai 1.8.0

Useagents in market2

Ringg AIPlatformfunding2026-08-26

Ringg AI extends Series A to $15M led by Peak XV

Funds an India-built multi-channel agent platform already in production at Flipkart/CRED-class operators — extension capital for WhatsApp and browser workflows, not a new speech model.

Also Connect

Details3 sources

Ringg AI closed a Series A extension that brings the round to $15 million, led by Peak XV Partners with Arkam Ventures and Capital 2B, to expand voice, WhatsApp, and browser agents beyond phone calls.

Ringg AI platform

GoogleModel providerlaunch2026-08-26

Gemini Live adds voice-driven Spark tasks, Daily Brief, and Gmail triage

Voice can start multi-step Gemini jobs across Docs/Sheets/Drive (AI Pro+) and a spoken inbox (AI Plus+), on the same day as Gemini 3.5 Transcribe.

Also Think

Details1 source

Google expanded Gemini Live with voice-driven multi-step Spark tasks, a spoken Daily Brief, hands-free Gmail triage, and Personal Intelligence memory.

Gemini Live