Hearspeech-to-text8
Cactus Compute ships Whistle, a 16.9 MB on-device ASR model
A sub-20 MB CPU ASR that shares Needle's tool-calling binary is a different on-device path from Fermion Phonon-2 (`2026-09-29-fermion-phonon-2`, 164 MB) — vendor WER claims still need independent checks beyond the blog tables.
Details4 sources · 3 X posts
Whistle is an open speech recognition model that runs on CPU in the same C++ engine as Needle, covering seven languages in a 16.9 MB file with vendor-claimed 11 ms time to first token.
Whistle
Twilio Batch Transcription Configuration reaches general availability
Post-call STT becomes a named, reusable config object with Deepgram Nova-3/Nova-2 or Twilio-managed engines — no redeploy to swap model or language.
Also Connect
Details1 source
Twilio made Batch Transcription Configuration generally available so recorded calls can be transcribed after completion via a reusable config for engine, speech model, language, and destination.
Batch Transcription Configuration
MicrosoftModel providermodel release2026-10-01
Microsoft ships MAI-Transcribe-2-Streaming at $0.54/hr, AA #1
Microsoft’s first dedicated streaming SKU on the Transcribe-2 line posts AA-WER Streaming #1 at $0.54/hr — the realtime counterpart to September’s $0.10/hr batch model.
Details3 sources · 2 X posts
Microsoft AI released MAI-Transcribe-2-Streaming, a low-latency streaming STT sibling to batch MAI-Transcribe-2, covering 60 languages with continuous language detection. Introductory price $0.54 per audio hour through year-end; public preview on Azure Speech / Foundry.
MAI-Transcribe-2-Streaming
Introducing 3 new models: MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash.
Accurate streaming transcription. Natural speech and less waiting between turns.
Build voice agents that keep the conversation moving!
https://x.com/MicrosoftAI/status/2105693024013467905
Microsoft AI has released MAI-Transcribe-2-Streaming, taking the #1 spot for Final Transcript accuracy and First Partial Transcript accuracy on AA-WER Streaming with 2.5% WER at 0.13s after end of speech
MAI-Transcribe-2-Streaming is @MicrosoftAI's new streaming Speech to Text
https://x.com/ArtificialAnlys/status/2105694108736188894
Devnagri ships Black Bird ASR, claiming lead on Voice of India telephonic bench
Vendor-claimed Indian telephonic ASR lead versus Sarvam and Gemini on an independent closed set — self-run API bake-off, not a third-party leaderboard publish — after the March Speech AI platform launch (`2026-03-09-devnagri-speech-ai`).
Details3 sources
Devnagri AI announced Black Bird, a new Indian-language ASR model, reporting 8.7 average OI-WER on the Voice of India telephonic benchmark across 15 languages — ahead of Sarvam Saaras V3 and Google Gemini 3 Pro on Devnagri’s own production-API bake-off.
Black Bird
Fireflies launches Talk, free unlimited voice dictation in its desktop app
A meeting-notes vendor undercuts $12–15/month standalone dictation by bundling unlimited desktop STT on Free — same company that shipped Voice Agents on 2 September (`2026-09-02-fireflies-voice-agents`).
Also Use
Details4 sources · 2 X posts
Fireflies shipped Talk inside its Mac and Windows desktop app: hold a hotkey, speak, and cleaned text lands at the cursor in any app, free and unlimited on every plan including Free, in more than 90 languages.
Fireflies Talk
FermionModel providermodel release2026-09-29
Fermion ships Phonon-2 open-weight English ASR at 164 MB
Shrinks open English ASR to a 164 MB CC-BY-4.0 download that Fermion says beats Whisper large on average WER and runs ~174× realtime on an M5 MacBook Air — vendor Open ASR Leaderboard scoring, one month after Phonon-1 (`2026-08-28-fermion-phonon-1`).
Details3 sources · 1 X post
Fermion released Phonon-2, a 164 MB English speech-recognition model that claims higher average accuracy than OpenAI Whisper large despite being about a tenth the download size, and transcribes about an hour of audio in roughly 20 seconds on a MacBook Air under CC-BY-4.0.
Phonon-2
Introducing Phonon-2: a new standard in speech recognition per byte. At just a tiny 164MB download, it is more accurate on average than OpenAI's Whisper large (a model 10X its size). Transcribe an hour of audio in just 20 seconds on a MacBook Air. Open weights, CC-BY-4.0.
https://x.com/yoitsmanan/status/2104990913886031993
AssemblyAI ships Universal-3.6 Pro Realtime for noisy agent calls
Short-response errors fall 45% on AssemblyAI's own 12,460-scenario agent bench (1.45% versus 2.65%) at held 3.5 Pro latency and $0.45/hr — the gain that matters on yes/no turns over noisy phone lines, with the benches still vendor-run rather than an independent audit.
Details1 source
AssemblyAI released Universal-3.6 Pro Realtime streaming STT for voice-agent and telephony audio, covering 32 languages at the same latency and $0.45/hr rate as Universal-3.5 Pro under model id universal-3-6-pro.
Universal-3.6 Pro Realtime
ModulateModel providerfunding2026-09-28net-new
Modulate raises $25M led by Future Ventures to scale audio-native Velma
A dated company primary for the 2026 Modulate raise that an earlier sweep skipped as funding-only — Velma sits beside the talking stack as an understanding and supervision layer.
Details2 sources · 1 X post
Modulate announced $25 million in new funding led by Future Ventures, with Hyperplane and Lakestar participating, bringing lifetime funding stated in the release to $60 million for its Velma audio-understanding platform.
Velma
Speaktext-to-speech12
Smallest AI partners with Tenstorrent to run Lightning V2 TTS on-prem
Gives Lightning V2 a dated on-prem Blackhole path with usage-based pricing — useful when GPU-class spend or data residency blocks cloud TTS for bank, telco, or healthcare agents.
Details2 sources · 2 X posts
Smallest AI and Tenstorrent opened a production on-prem path that runs Lightning V2 realtime TTS natively on Tenstorrent Galaxy Blackhole servers, with usage-based pricing and no upfront hardware commitment.
Lightning V2 on Tenstorrent · vendor Tenstorrent
Index TeamModel providermodel release2026-10-02net-new
Index Team releases Index-Echo-S2ST-9B for voice-preserving dubbing
Open dubbing weights with explicit source-voice conditioning, which is speak plus a speech-translation front end, not a conversational agent that decides what to say.
Also Hear
Details2 sources · 1 X post
Index-Echo-S2ST-9B is an Apache-2.0 speech-to-speech dubbing package that takes Chinese or English speech and speaks a translation in English, Spanish, Japanese, or Chinese while conditioning on the source speaker.
Index-Echo-S2ST-9B
speak once, hear yourself in another language 🗣️🌍
Index-Echo-S2ST is a new open model that dubs your speech and keeps your voice: talk in English or Chinese, get it back in English, Chinese, Japanese or Spanish
▶️ on Spaces
https://x.com/HuggingApps/status/2106003866974065097
SunoPlatformlaunch2026-10-01net-new
Suno opens Speech (beta), spoken audio with matching background music
Brings a consumer music product into the speak layer with one model that writes voice and score together, rather than a developer TTS API. Other 2026 Suno music drops stay out of this filing.
Details2 sources · 1 X post
Speech (beta) generates spoken audio and original background music as one track inside Suno on web and mobile, from an idea, a poem, or text you already wrote plus a voice and musical style.
Speech (beta)
MicrosoftModel providermodel release2026-10-01
Microsoft ships MAI-Voice-2.1-Flash TTS at $15/1M characters, 150 ms e2e
Pairs with same-day Transcribe-2-Streaming (`2026-10-01-microsoft-mai-transcribe-2-streaming`) so Microsoft’s hear and speak SKUs share one Foundry/Vercel path at agent-grade latency.
Details4 sources · 1 X post
Microsoft AI released MAI-Voice-2.1-Flash, a low-latency sibling to MAI-Voice-2.1 for high-volume voice agents, with the same 23-language coverage and ~150 ms end-to-end latency for up to 45 seconds of audio.
MAI-Voice-2.1-Flash
Introducing 3 new models: MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash.
Accurate streaming transcription. Natural speech and less waiting between turns.
Build voice agents that keep the conversation moving!
https://x.com/MicrosoftAI/status/2105693024013467905
MicrosoftModel providermodel release2026-10-01
Microsoft ships MAI-Voice-2.1 multilingual TTS at $22/1M characters
Cross-language speaker consistency at 23 languages / 26 locales lets one brand voice stay native across English, Mandarin, German, and peers without swapping voices mid-session.
Details4 sources · 1 X post
Microsoft AI released MAI-Voice-2.1, its strongest multilingual text-to-speech model yet, covering 23 languages and 26 locales with a single speaker identity that keeps a native accent when switching languages.
MAI-Voice-2.1
Introducing 3 new models: MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash.
Accurate streaming transcription. Natural speech and less waiting between turns.
Build voice agents that keep the conversation moving!
https://x.com/MicrosoftAI/status/2105693024013467905
LiveKit Agents 1.8.4 adds Microsoft AI TTS, Eleven v4, and SmallestAI continuations
Wires same-day MAI-Voice-2.1/Flash and Eleven v4 into Agents while SmallestAI gains TTS continuations — builder path without waiting for a minor.
Details1 source
livekit-agents 1.8.4 adds a Microsoft AI speech plugin, ElevenLabs eleven_v4 and eleven_v4_turbo, SmallestAI TTS continuations protocol support, and a Deepgram TTS speed parameter.
livekit-agents 1.8.4
GradiumModel providermodel release2026-10-01
Gradium makes its sub-50 ms TTFA TTS the default after beta
Moves the under-50 ms TTFA path from `gradium-tts-beta` opt-in to the production default — Speko naturalness claim is vendor-cited, not re-run here.
Details2 sources · 1 X post
Gradium promoted its low-latency Text-to-Speech model out of beta as the default API and Studio path, stating around 50 ms time to first audio with hard-case pronunciation retained.
Gradium TTS
Our default Text-to-Speech model now has around 50ms time to first audio.
The model scores highest on naturalness for any sub-100ms model on Speko public benchmarks.
How much of your turn budget does TTS take?
https://x.com/GradiumAI/status/2105689201278664804
Deepgram self-hosted 261001 watermarks Flux TTS and drops Whisper
Self-hosted Flux TTS operators hit a hard watermarker requirement on 1 October while Whisper is removed as a supported path, so on-prem speak stacks need a coordinated Engine and Helm change rather than a quiet container bump.
Also Hear
Details1 source
Deepgram’s October 2026 self-hosted release 261001 requires Flux TTS watermarking, removes Whisper model support, aligns FIPS Engine metrics with standard images, and improves Japanese and German formatting.
Self-Hosted release 261001
Deepgram adds Flux TTS inline pause and IPA pronunciation controls
Drug names and intentional silence land as first-class Flux TTS markup instead of Aura-style ellipsis prompting — pronunciation is still Early Access and cannot stack with pause or non-default speed.
Details3 sources
Deepgram shipped escaped pause markers and Early Access inline IPA pronunciation on Flux TTS (/v2/speak), alongside the existing speed control, with combination guardrails that reject degrading mixes.
Flux TTS inline pause and pronunciation
Telnyx Ultra adds seven Urdu and eight Odia TTS voices
South Asian Ultra coverage lands three days before the Chennai SIP anchorsite (`2026-09-30-telnyx-chennai-sip-anchorsite`), so India/Pakistan language TTS and India media PoPs arrive in the same W40 CPaaS push.
Also Connect
Details1 source
Telnyx Ultra now synthesises Urdu and Odia natively with fifteen new catalog voices usable from Voice AI Assistants, the TTS API, Call Control speak, and TeXML without a language hint.
Telnyx Ultra Urdu and Odia TTS voices
KyutaiModel providerother2026-09-28
Kyutai publishes Pocket TTS training via drifting one-step objective
Documents a Jacobian-free one-step recipe for the same Pocket TTS continuous-latent stack filed in W03 (`2026-01-13-kyutai-pocket-tts`) — vendor <1% WER / cloning claims, and an english_drifting_26-09 checkpoint path, without restating the January model launch.
Details4 sources · 1 X post
Kyutai published a technical blog on training the Pocket TTS sampler head with drifting, a one-step generative objective from Deng et al., claiming under 1% WER with voice cloning and calling it the first speech and first autoregressive model trained that way.
Pocket TTS drifting training
ElevenLabs ships Eleven v4 and low-latency v4 Turbo for expressive TTS
Puts a #1-ranked expressive TTS path and a ~100 ms Turbo SKU on the same architecture for ElevenAgents builders — vendor preference and latency figures are ElevenLabs' September 2026 benches, not an independent audit.
Details4 sources · 1 X post
ElevenLabs released Eleven v4 and Eleven v4 Turbo text-to-speech models with 90+ languages, richer inline delivery tags, and Instant Voice Clones from about 10 seconds of audio, with Turbo aimed at realtime voice agents.
Eleven v4 and Eleven v4 Turbo
Orchestrateframeworks, evals7
OnepinPlatformlaunch2026-10-01net-new
Onepin launches a TTS production layer for pronunciation and line QC
Production TTS failures are often brand-name and number misreads, not model naturalness — Onepin sells a shared pronunciation and line-QC layer across 30+ engines so teams stop regenerating whole takes for one wrong word.
Also Speak
Details2 sources · 2 X posts
Onepin launched a production step after text-to-speech that checks every voiceover line against a 4-million-word pronunciation dictionary, normalises prices and dates, scores naturalness, clarity, and word accuracy, and fixes a single wrong word in the same voice without re-rendering the take.
Onepin
Decagon launches Duet Apprentice beta for learning from docs and Slack
Agent builders still paste SOPs by hand while policies drift in Slack — Apprentice closes that loop on the same Dialogues day as Voice 3, with a vendor claim that Duet already drafts more than 70% of live AOPs.
Details2 sources · 1 X post
Decagon introduced Duet Apprentice in beta so Duet can onboard from Notion, Confluence, SharePoint, or Google Drive, learn from escalated human-rep conversations, and keep updating Agent Operating Procedures from Slack and Microsoft Teams activity.
Duet Apprentice
Decagon launches Agent Modules for journeys beyond support
Concierge platforms that stop at ticket deflection leave lead-to-collections orchestration to custom builds — Modules package those journeys with Simulations, Campaign Composer, and shared memory on the same 1 October Dialogues stack as Voice 3.
Also Use
Details2 sources · 1 X post
Decagon introduced Agent Modules, journey-specific infrastructure, analytics, and Campaign Composer for lead qualification, onboarding, collections, and growth use cases that reuse one agent with shared memory across instances.
Agent Modules
Hume AIModel providerother2026-09-29
Hume publishes a five-dimension multi-speaker conversational TTS evaluation
Same-gender speaker separation remains the failure mode overall scores hide — Gemini mixed-gender pairs average 4.65 while male-male averages 2.87, and ElevenLabs v3 flips the pattern (male-male 4.33, female-female 3.02) — so dialogue TTS buyers need this rubric beside cloning and controllability boards (`2026-09-10-hume-ai-voice-replication-leaderboard`, `2026-09-17-hume-voice-controllability-leaderboard`).
Also Speak
Details2 sources · 1 X post
Hume published a bespoke multi-speaker dialogue rubric with 48 two-speaker scripts and blind human ratings across five dimensions, applying it to Gemini TTS generations and peer multi-speaker systems.
Multi-speaker TTS evaluation
SierraPlatformlaunch2026-09-28
Sierra turns Ghostwriter into a proactive Slack and Teams teammate for agent ops
Agent improvement loops that used to wait for a weekly call review now start from an always-on teammate that reads live voice and chat traffic — useful for Sierra operators who already run production agents, with the caveat that the post is still a controlled rollout rather than a numbered GA date.
Also Use
Details1 source
Sierra released a Ghostwriter update that lives in Slack and Teams, reviews customer-agent conversations, and proposes experiments without waiting for a prompt.
Ghostwriter
OcularPlatformlaunch2026-09-28
Ocular launches a full-duplex audiovisual conversational dataset
Video conversational and interaction models need four streams on one clock — speaker A/B audio plus video — and public full-duplex A/V sets remain thin (ViCo 1.6 h; ViCo-X 0.4 h in the INFP survey), so a 1,000+ hour on-request corpus is the video counterpart to Ocular's earlier speech Hi-Fi work.
Also HearSee
Details2 sources · 1 X post
Ocular released a full-duplex audiovisual dataset of two-person video calls with four synchronised streams — per-speaker audio and video on one shared clock — built for video conversational models.
Full-Duplex Audiovisual Dataset
CekuraPlatformlaunch2026-09-28
Cekura opens a speech-to-speech phone-agent benchmark across nine realtime models
Long-form Medicare reliability separates the field far more than short clinic bookings — GPT Realtime 2.1 holds 61% pass³ on the long form while several peers fall under 15%, so teams picking a duplex model for multi-field phone intake get a shared, public call log instead of vendor demos alone.
Also Think
Details2 sources · 1 X post
Cekura published an open speech-to-speech benchmark that runs the same Pipecat phone agent on live calls across nine realtime voice models over 82 scenarios and three runs each.
Speech-to-speech benchmark
Useagents in market9
DecagonPlatformpartnership2026-10-02
Decagon joins OpenAI Marketplace as a launch partner for CX agents
OpenAI commitment spend as a buy path for Decagon agents — distribution, not a new voice SKU, three days after Voice 3/Chord at Dialogues.
Also Orchestrate
Details2 sources · 2 X posts
Decagon became a launch partner on OpenAI's B2B marketplace so eligible enterprise customers can apply existing OpenAI commitments toward purchasing Decagon customer-facing agents.
OpenAI Marketplace · vendor OpenAI
PayNearMe launches an AI Servicing and Collections Agent in PayXM
Fintech-embedded outbound voice for collections/servicing — pilot throughput (~28 staff-hours of dials in one hour) is vendor-claimed from Indiana Finance Company, not an independent audit.
Also Connect
Details3 sources · 1 X post
PayNearMe shipped a voice-and-text agent inside PayXM that places outbound calls, answers inbound, and helps customers make or schedule payments, escalating to staff when needed.
AI Servicing and Collections Agent
Today, we're launching the AI Servicing and Collections Agent, a new agent built directly into the PayXM™ platform to expand inbound self-service and outbound engagement.
Read the announcement: https://home.paynearme.com/press/paynearme-launches-ai-servicing-and-collections-agent-to-help-support-teams-scale-while-lowering-the-cost-to-serve/
https://x.com/PayNearMe/status/2105634124904227324
VapiPlatformlaunch2026-09-30
Vapi launches Personal Agent Calling so consumer assistants can place phone calls
Consumer-agent telephony moves from Muse/Instinct-class product features into a shared voice layer underneath many assistants, with a dated beta and a worked 14-minute hold example on 30 September.
Also Connect
Details1 source
Vapi opened Personal Agent Calling in beta at phone.vapi.ai, letting personal assistants such as Muse, Instinct, ChatGPT Work and Claude Code dial businesses, hold the conversation and report results on the user’s Vapi minutes.
Personal Agent Calling
RiachueloEnterprisepartnership2026-09-30net-new
Riachuelo runs Genesys Agentic Virtual Agent on WhatsApp for Brazilian retail CX
Puts a major Brazilian retailer on Genesys AVA with an 84% vs 30% self-service jump on WhatsApp — still a digital channel go-live, with voice journeys only described as future exploration.
Also Orchestrate
Details3 sources · 1 X post
Genesys said Brazilian fashion retailer Riachuelo is using Genesys Cloud Agentic Virtual Agent on WhatsApp Business so shoppers can resolve delivery and support issues without a human, after earlier Genesys Cloud Copilot for employees.
Genesys Cloud Agentic Virtual Agent (WhatsApp) · vendor Genesys
Marriott expands Cresta AI Agent for guest voice and chat answers around the clock
Puts a major hotel group’s guest-facing voice+chat path onto Cresta AI Agent as an expansion of an existing relationship — still Cresta’s X primary without a Marriott wire, so scope and metrics stay thin.
Also Connect
Details2 sources · 1 X post
Cresta said Marriott International is expanding their relationship so Cresta AI Agent handles guest inquiries across voice and chat inside Marriott’s patent-pending tech platform, giving faster multilingual answers while associates keep higher-touch work.
Cresta AI Agent (guest experience expansion) · vendor Cresta
SierraPlatformpartnership2026-09-29
Sierra joins OpenAI Marketplace as a launch partner for enterprise agents
Distribution via OpenAI spend commitments, not a new voice SKU — sits beside same-week Ghostwriter teammate (`2026-09-28-sierra-ghostwriter-teammate`) and W39 Liberty Global×Sierra (`2026-09-23-liberty-global-sierra-partnership`).
Also Orchestrate
Details2 sources · 1 X post
Sierra became a launch partner in OpenAI’s enterprise Marketplace so eligible customers can build Sierra agents and apply existing OpenAI commitments toward the purchase.
OpenAI Marketplace · vendor OpenAI
Proud to be a launch partner in @OpenAI's new marketplace! Now, eligible OpenAI customers can start building with Sierra, using existing OpenAI commitments.
We're excited to deepen our partnership with OpenAI, and make it even easier for businesses to build better customer
https://x.com/SierraPlatform/status/2104993581400514967
ElevenLabs lists ElevenAgents on the OpenAI Marketplace as a launch partner
Same DevDay distribution channel as Sierra (`2026-09-29-sierra-openai-marketplace`) for a TTS/agent vendor that shipped Eleven v4 the day before (`2026-09-28-elevenlabs-v4`).
Also SpeakOrchestrate
Details2 sources · 1 X post
ElevenLabs joined OpenAI’s enterprise Marketplace as a launch partner so eligible customers can access ElevenAgents and apply OpenAI commitments toward purchases once buying opens.
ElevenAgents OpenAI Marketplace · vendor OpenAI
ElevenAgents is now on the OpenAI Marketplace.
If your company is an OpenAI enterprise customer, you'll soon be able to give it a voice in 90+ languages and put it on your existing OpenAI commitment.
Learn more: https://elevenlabs.io/blog/elevenagents-now-on-the-openai-marketplace
https://x.com/ElevenLabs/status/2105004567335420258
Rabona AI opens a blended AI call centre that switches inbound and outbound on one agent
Japan-market inbound/outbound blend with named go-lives — Houmiya at ~90% AI-completed post-install inbound and ZAP at ~80% fewer delivery-date calls — plus an outcome-priced SKU for some accounts from 28 September.
Also Connect
Details3 sources
Tokyo University spinout Rabona AI began offering a blended AI call centre in which one agent handles inbound and outbound phone work with CRM context, post-call write-back, and SMS, citing Houmiya Equipment and ZAP as early operators.
Blended AI Call Centre
PrestoPlatformfunding2026-09-28
Presto raises $10M from Remus Capital affiliates to scale QSR Voice AI
A dated primary for a QSR drive-thru vertical already live with national brands, with Toast and ElevenLabs named in the same release window rather than a greenfield demo raise.
Details2 sources
Presto Phoenix announced $10 million from Remus Capital-affiliated investors and other existing backers to expand Presto Voice drive-thru deployments, product work, and QSR customer impact.
Presto Voice