Microsoft launched MAI-Transcribe-1.5 at its Build 2026 developer conference on June 2. It is a speech-to-text tool built entirely in-house, not sourced from a third party, and it is already embedded in the products that power most Microsoft-based business operations: Teams, Copilot, Dynamics 365 Contact Center, and GitHub.
The clearest upgrade is speed. The tool can transcribe one full hour of recorded audio in under 15 seconds. The previous version took around 53 seconds for the same job. For a business that processes hundreds of customer calls or meeting recordings daily, this is not a minor improvement. It means workflows that previously needed overnight batch processing can now run in near real-time.
Language coverage grew from 25 to 43. Ten of the 18 new languages are South Asian, including Bengali, Tamil, and Telugu. Eight are European, including Ukrainian, Greek, and Catalan. For any business operating across multiple countries, that broadening matters, especially in markets where the previous generation simply did not work well enough to rely on.
The keyword biasing feature addresses a long-standing frustration with general-purpose transcription. Generic tools regularly mishear names, internal product codes, medical terms, and industry-specific abbreviations. With MAI-Transcribe-1.5, you can supply up to 200 specific words, and the model will factor them into its predictions rather than defaulting to the nearest common word. Microsoft reports a 30% drop in transcription errors on specialist vocabulary when this feature is active. That is particularly relevant for insurance, healthcare, legal, and any sector where precise terminology in a recorded conversation carries compliance or liability weight.
On accuracy benchmarks, the picture is mixed. Microsoft claims the top position on FLEURS, a standard multilingual transcription test covering 43 languages, outperforming tools from OpenAI, Google, and ElevenLabs. On the independent Artificial Analysis leaderboard, it sits third overall with a 2.4% error rate. Both results are credible for a production-grade tool, and third-party verification typically follows within four to eight weeks of launch.
Pricing is confirmed at $0.36 per hour of audio, unchanged from the previous version. That is competitive with comparable tools from Deepgram and AssemblyAI.
The broader context is worth noting. The global speech-to-text market was valued at around $3.3 billion in 2025 and is growing quickly. Meeting transcription in particular is the fastest-growing segment, driven by the normalisation of hybrid work and the volume of recorded conversations organisations now generate daily.
For businesses already running on Microsoft infrastructure, the practical takeaway is simple: better transcription is arriving inside tools you already pay for. The keyword biasing feature, in particular, is worth testing if your team relies on accurate capture of specialised language in recorded conversations. You do not need to adopt a new tool or sign a new contract. It is already there.