Microsoft releases three models for speech transcription and generation for voice agents
Microsoft has released MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. They offer transcription in 60 languages and speech generation in 23 languages. The introductory transcription price is 0.54 USD per hour of audio until the end of the year.
Microsoft has released MAI-Transcribe-2-Streaming for real-time speech transcription and MAI-Voice-2.1 and MAI-Voice-2.1-Flash for speech generation. All are available through Microsoft Foundry and MAI Playground; both speech generation models are also available through OpenRouter.
MAI-Transcribe-2-Streaming transcribes 60 languages. According to Microsoft, it delivers the first partial results in just over 100 milliseconds and ranks first for accuracy on Artificial Analysis. The company says this allows voice agents to respond while the user is still speaking. The introductory price until the end of the year is 0.54 USD per hour of audio.
MAI-Voice-2.1 supports speech generation in 23 languages using the same voice and, according to Microsoft, with a native accent in each language. According to the company, MAI-Voice-2.1-Flash has a latency of 150 milliseconds and a price of 15 USD per million characters instead of the quoted 22 USD per million characters. Both voice models allow voice cloning from a few seconds of reference audio and include safeguards intended to prevent misuse. In one test, approximately half of the 4 000 participants believed the generated voices were those of a real person.
Why it matters
Voice agent developers can combine streaming transcription in 60 languages with speech generation. The reported latencies matter for the length of pauses in a conversation, but actual performance needs to be verified in the specific application. Prices per hour of audio and per million characters provide a basis for estimating operating costs. Voice cloning and support for multiple languages may help recording creators prepare versions in different languages.
Two audiences, two different impacts
What this means
For individuals
A voice recording creator can try cloning their own voice from a short sample and assess the use of MAI-Voice-2.1 for recordings in multiple languages.
For a business
Companies developing voice agents gain another option for transcription and speech generation. When designing the conversation flow, they can assess the reported latencies and calculate costs based on the volume of audio and generated characters; the introductory transcription price applies only until the end of the year.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.