hotwirednews Install bot

Week 31

27 Jul – 2 Aug 20262026-W31 · 13 filings · 9 companies · 4 net-new

The Hear layer led the week with 5 of 13 filings, from Daily / Pipecat, Gladia, Edge0, OpenAI and 1 more.

Hearspeech-to-text5

Daily / PipecatPlatformAPI2026-08-01

Pipecat 1.7.0 adds STT usage metrics and ElevenLabs background-audio filter

Usage metering and cleaner Scribe transcripts matter for production cost and IVR noise.

Details3 sources

pipecat-ai 1.7.0 expands speech-to-text usage metrics and adds filter_background_audio on ElevenLabsRealtimeSTTService, plus shared http_client for OpenAI TTS and Whisper STT.

pipecat-ai 1.7.0

GladiaModel provideracquisition2026-07-31net-new

OVH Groupe completes acquisition of Gladia speech-to-text

A European sovereign-cloud operator internalises a multilingual STT API rather than renting US speech models — Gladia stays branded, but compute and residency now sit on OVH.

Details3 sources

OVH Groupe announced it completed the acquisition of Paris speech-to-text company Gladia, following exclusive negotiations disclosed on 11 June 2026.

Gladia STT

Edge0Model providermodel release2026-07-30net-new

Edge0 open-sources ARK-ASR-3B multilingual speech recognition

Largest Edge0 ASR open weight before Audio8-ASR-0.1B and Infinite — vendor Open ASR English average WER 5.04% on the 3B card, Apache-2.0, nineteen languages.

Details5 sources · 1 X post

Edge0 released ARK-ASR-3B, an Apache-2.0 multilingual ASR model claiming a 5.04% average English WER on the Hugging Face Open ASR Leaderboard short-form suite across nineteen languages, alongside a smaller ARK-ASR-0.6B sibling.

ARK-ASR-3B

ARK-ASR-3B: #1 overall; ARK-ASR-0.6B: #5 overall. Both support 19 languages and are released under Apache 2.0.

https://x.com/SamuelZengML/status/2082800506003767590
OpenAIModel providermodel release2026-07-29

GPT-Transcribe and GPT-Live-Transcribe: context-aware ASR

Promptable ASR. Community cite for the Live variant: context-aware semantic accuracy 38.5% to 44.6% with free-form context.

Details2 sources

OpenAI introduced two context-aware transcription models, streaming GPT-Live-Transcribe and async GPT-Transcribe, that take prompts, keywords, and language hints.

GPT-Transcribe

NablaPlatformlaunch2026-07-28net-new

Nabla launches on-device medical dictation for Mac with Epic

On-device clinical dictation on Mac with Epic — privacy and latency pitch versus cloud STT, bundled with ambient rather than a separate vendor.

Also Use

Details3 sources

Nabla shipped Nabla Dictation for Mac, a fully on-device clinical dictation app built with Core ML and Apple silicon that pairs with Nabla ambient documentation and Epic workflows without sending audio to the cloud.

Nabla Dictation for Mac

Speaktext-to-speech3

Daily / PipecatPlatformAPI2026-08-01

Pipecat 1.7.0 adds Azure TTS force_locale

Forces Azure locale when auto language detection would pick the wrong voice pack.

Details3 sources

pipecat-ai 1.7.0 adds opt-in force_locale on AzureTTSService and AzureHttpTTSService, and reach_inactive_services on settings update frames.

pipecat-ai 1.7.0

Edge0Model providermodel release2026-07-29net-new

Edge0 open-sources Audio8 TTS Preview 0.6B with zero-shot cloning

Compact DualAR zero-shot TTS under Apache 2.0 with a bundled 44.1 kHz codec — vendor Seed-TTS EN WER 1.506 / ZH CER 0.950 on the launch thread — and the speak-layer sibling to Edge0's later ASR Infinite filing (`2026-09-23-edge0-audio8-asr-infinite`).

Details5 sources · 1 X post

Edge0 released Audio8 TTS Preview 0.6B, an Apache-2.0 multilingual text-to-speech model with zero-shot voice cloning, a DualAR architecture, and a bundled 44.1 kHz neural codec across eleven recommended languages.

Audio8 TTS Preview 0.6B

Today, we're open-sourcing Audio8-TTS Preview: 11 languages. Zero-shot voice cloning. A bundled 44.1 kHz codec. Apache 2.0.

https://x.com/SamuelZengML/status/2082430003460166142
Fish AudioModel providerfunding2026-07-28

Fish Audio raises $52M seed led by Coreline Ventures and Capital Today

Dates the seed behind the already-filed S2 open-source and S2.1 Pro free-API ships — a creator-to-enterprise TTS lab with a named LiveKit/Retell integration plan, not a new model drop.

Details2 sources

Fish Audio announced $52 million in seed funding led by Coreline Ventures and Capital Today as it marked its first anniversary, citing $21 million ARR and more than 8 million users.

Fish Audio

Showavatars, video1

Daily / PipecatPlatformAPI2026-08-01

Pipecat 1.7.0 adds LemonSlice local background images and faster Tavus audio

Avatar sessions start faster and look branded without remote background downloads.

Details3 sources

pipecat-ai 1.7.0 lets LemonSlice transport take a local image as the avatar background and defaults TavusParams.audio_out_faster_than_realtime to True.

pipecat-ai 1.7.0

ThinkLLMs, speech-to-speech2

PolyAIModel providermodel release2026-07-30net-new

PolyAI Dialog-RSN-1: audio-native dialog, TTS left to a separate model

Contact-centre stacks that need a branded TTS can take audio-aware input without moving the call onto a speech-to-speech model.

Also Hear

Details1 source · 1 X post

PolyAI published Dialog-RSN-1, a request-based model that does turn-taking, ASR, function calling, and response, and leaves speech generation to a separate TTS.

Dialog-RSN-1

The traditional cascade loses the audio. Speech-to-speech loses control. Dialog-RSN-1 loses neither. We're introducing an entirely new approach to building voice agents, one that hears every caller's audio directly.

https://x.com/polyaivoice/status/2082819930383130879
xAIModel providermodel release2026-07-29

Grok Voice Think Fast 2.0: speech-to-speech at $0.08/min

Full-duplex STS at $0.08/min with an A/B test on the Starlink support line. Became grok-voice-latest on 5 Aug.

Details2 sources

xAI released Grok Voice Think Fast 2.0, a speech-to-speech model with 82.9% on AA STS, 0.70s time-to-first-audio, and $0.08/min pricing.

Grok Voice Think Fast 2.0

Connecttelephony, channels1

Daily / PipecatPlatformAPI2026-08-01

Pipecat 1.7.0 adds LiveKit inbound SIP DTMF

LiveKit SIP calls can collect digits the same way Daily and RTVI already could.

Details3 sources

pipecat-ai 1.7.0 adds inbound SIP DTMF on LiveKitTransport via on_dtmf_event and InputDTMFFrame.

pipecat-ai 1.7.0

Useagents in market1

InfinitusPlatformlaunch2026-07-30

Infinitus launches AI-first hub for specialty pharma patient support

Packages Infinitus voice agents into a specialty-pharma hub SKU billed on outcomes and work performed, competing with legacy hub call centres rather than only selling payor diallers.

Also ThinkConnect

Details3 sources

Infinitus launched an AI-first, human-backed hub that uses voice and digital AI agents for specialty-therapy enrollment, benefits verification, prior authorisation, adherence, and after-hours support, with humans kept on judgment calls.

AI-first hub solution