Soniox ships v5 Real-Time STT with live speaker separation
Pairs the structured-speech v5 stack with a live path five days after Async — endpoint_sensitivity and speaker-aware streaming matter for agent turn-taking, not just WER.
Also Orchestrate
Details2 sources
Soniox released stt-rt-v5 for live audio, adding reinvented real-time speaker separation, spoken language ID across 60+ languages, streaming translation over 3,600 pairs, faster semantic endpointing with endpoint_sensitivity, and stronger alphanumeric formatting.
Soniox v5 Real-Time (stt-rt-v5)