Microsoft's New Speech AI Transcribes Before You Finish Speaking in 100ms
Microsoft released MAI-Transcribe-2-Streaming for real-time transcription with 100ms latency, plus MAI-Voice-2.1 and Flash for synthesis; the API streams revisable intermediate results, requires apps to manage commit timing, supports 60 languages, and is in public preview on Azure Foundry.
On October 1, Microsoft announced MAI-Transcribe-2-Streaming , a streaming speech-to-text model that returns preliminary transcripts within 100 milliseconds of receiving audio — before the speaker has finished talking. Alongside it, Microsoft released two speech synthesis models: MAI-Voice-2.1 and MAI-Voice-2.1-Flash . All three models are available in public preview on the Azure AI Foundry platform but are not yet considered production-ready.
Streaming Transcription vs. Batch Recognition
Unlike traditional batch transcription that processes an entire recording at once, MAI-Transcribe-2-Streaming emits results incrementally as audio arrives. This enables live captioning and real-time information lookup while the user is still speaking. The model supports 60 languages and can automatically detect the spoken language continuously during a session.
Intermediate and Delta Results
The API distinguishes two types of incremental output:
delta — confirmed segments that are appended to the final transcript and will not change.
intermediate — provisional results that replace all unconfirmed portions whenever new audio is processed.
Simply concatenating every response leads to duplication because intermediate results overwrite earlier provisional text. Developers must handle these two streams separately.
Design Implications for Voice Agents
The article illustrates the practical impact with a booking scenario: a user says "Next Friday, or rather Thursday." The initial transcript "Friday" could trigger a room-availability search, but if the system commits the booking before the correction to "Thursday" arrives, an error occurs. The recommended pattern is to separate reversible actions (search, suggestion retrieval) that can use intermediate results from irreversible actions (confirmation, payment) that must wait for delta (confirmed) segments.
Endpoint Detection and Session Limits
The real-time API does not automatically detect speech endpoints or send commit requests. The application must decide when a natural pause or the end of an utterance occurs and send a commit notification. Combining voice activity detection (VAD) to distinguish brief breaths from true utterance boundaries is critical: committing too early processes incomplete utterances; waiting too long negates the low-latency advantage. Each session is limited to one hour.
API Format and Comparison
The streaming API uses a message format similar to OpenAI's Realtime API, but it lacks MAI-specific features such as intermediate results and the explicit commit mechanism. A reference to the model catalog page is provided for further details.
Reference: https://ai.azure.com/catalog/models/MAI-Transcribe-2-Streaming?publisher=microsoft
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
21CTO
21CTO (21CTO.com) offers developers community, training, and services, making it your go‑to learning and service platform.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
