LiveKit Agents 1.4.0 adds Azure STT TrueText post-processing
STT-only surface for 1.4.0.
Details1 source
livekit-agents 1.4.0 adds a TrueText post-processing option on Azure STT.
livekit-agents 1.4.0
The Hear layer led the week with 4 of 12 filings, from LiveKit, Soniox, Mistral AI and DeepL.
STT-only surface for 1.4.0.
livekit-agents 1.4.0 adds a TrueText post-processing option on Azure STT.
livekit-agents 1.4.0
Streaming sibling to the 29 January v4 Async cut — semantic endpointing and manual finalisation target agent barge-in failures that silence-only VAD still causes.
Soniox released stt-rt-v4 for low-latency streaming transcription across 60+ languages, with millisecond finality, semantic endpointing, manual finalization, and simultaneous streaming speech translation.
Soniox v4 Real-Time (stt-rt-v4)
Open-weight Realtime STT under Apache 2.0 at claimed sub-200ms delay gives voice-agent builders a self-host path alongside a $0.003/min batch API — WER claims are Mistral's benches.
Mistral AI released Voxtral Transcribe 2 as Voxtral Mini Transcribe V2 for batch jobs and Voxtral Realtime for live transcription, with diarization, context biasing, word-level timestamps, and Apache 2.0 open weights for Realtime.
Voxtral Transcribe 2
Contact-centre builders get a streamed STT-plus-translation WebSocket on 2 February 2026 — months before DeepL's broader Voice-to-Voice suite in April.
Also Connect
DeepL opened the DeepL Voice API to API Pro customers so applications can stream audio and receive a source-language transcript plus translations into up to five target languages in real time.
DeepL Voice API
TTS providers clustered into one speak row.
livekit-agents 1.4.0 adds a Camb.ai TTS plugin among speak-path provider work.
livekit-agents 1.4.0
India-first TTS with 35+ voices and streaming for agents — Josh Talks telephony preference is third-party; full-band still trails ElevenLabs v3 alpha on the same study.
Sarvam AI released Bulbul V3, a production text-to-speech model with 35+ professional voices across 11 Indian languages, low-latency streaming, consent-based voice cloning, and code-switching for Hinglish and other mixed speech.
Bulbul V3
Avatar plugins filed as one show cluster. Class-2: folded quiet patch(es) 1.4.5 into this row.
livekit-agents 1.4.0 adds an Avatario avatar plugin and customisable Bithuman GPU avatar endpoint handling.
livekit-agents 1.4.0
LLM/Responses cluster separate from orchestration APIs.
livekit-agents 1.4.0 adds xAI Responses LLM and Azure OpenAI Responses integrations.
livekit-agents 1.4.0
Major orchestration API surface for the 1.4 line across Python and Node Agents. Class-2: folded quiet patch(es) 1.4.2, 1.4.5 into this row.
livekit-agents and @livekit/agents 1.4.0 add LLMStream.collect() for LLM use outside AgentSession, stable tool IDs, commit_user_turn for manual realtime turns, and AgentConfigUpdate with initial judges; the Node package lands on the same 1.4 feature line.
livekit-agents / @livekit/agents 1.4.0
Voice finally shares the same Coaching and Playbooks path as chat and email from early February — one reasoning layer instead of per-channel stacks.
Also UseConnect
Ada began rolling out a Unified Reasoning Engine that replaces channel-specific stacks with one shared intelligence layer across Voice, Messaging, and Email, bringing Coaching and Playbooks onto voice for the first time.
Unified Reasoning Engine
Shifts public LLM comparison toward the multi-turn, tool-heavy profile voice agents actually need — and shows accuracy leaders still miss the latency budget that production cascades require.
Daily released aiewf-eval, an open benchmark that scores text-mode and speech-to-speech LLMs on tool calling, instruction following, knowledge grounding, and latency over ~30-turn voice-agent sessions.
aiewf-eval LLM voice-agent benchmark
Native EHR ambient scribe from Epic itself — raises the bar for third-party ambient vendors that sell into Epic shops.
Also Hear
Epic released AI Charting, a built-in Art feature that listens during visits, drafts the clinician note, suggests orders from the conversation, and accepts voice commands to reshape note structure.
AI Charting (Art)