Microsoft launches MAI-Transcribe-2 for faster, cheaper speech recognition

Microsoft's MAI-Transcribe-2 enters public preview with 60-language support, diarization and vendor claims focused on accuracy, speed and cost.

Mason Reed

Microsoft has introduced MAI-Transcribe-2, a new speech-recognition model aimed at transcription workloads where accuracy, speed and price all matter. The company says the model improves on its earlier transcription work while adding practical features for meetings, captions, call analysis and other applications that need structured speech-to-text rather than a bare transcript.

Microsoft says MAI-Transcribe-2 covers 60 languages and includes speaker diarization, word-level timestamps, automatic language identification, code switching and keyword biasing for specialist terminology. It is available in public preview through Microsoft Foundry and Azure Speech.

Microsoft is making ambitious benchmark claims

The headline numbers come from Microsoft, so they need to be treated as vendor results rather than universal production guarantees. Microsoft reports an average word-error rate of 5.2% on the FLEURS benchmark across 60 languages and says the model sits on the accuracy-and-latency Pareto frontier measured by Artificial Analysis. The company also describes it as substantially faster than competing systems in its tests.

Those claims are useful signals, but speech recognition is unusually sensitive to the real material being transcribed. Microphone quality, accents, overlapping speakers, background noise, specialist names and the way an application chunks audio can all change results. A team choosing a transcription model still needs to test representative recordings rather than assuming a leaderboard order will survive every workload.

The operational features may matter more than one score

Diarization helps identify who spoke. Word-level timestamps make it easier to connect a transcript back to the source recording. Keyword biasing gives applications a way to steer recognition toward domain-specific names and terms. Automatic language detection and code switching matter when conversations move between languages instead of staying in one clean language bucket.

Microsoft is also competing on price. The company is promoting introductory pricing of $0.10 per audio hour through the end of 2026. That can make large transcription trials easier to justify, although the launch rate should not be mistaken for a permanent price promise beyond the stated period.

The practical takeaway is that Microsoft now has a first-party model positioned for production-style transcription rather than a research demo. Its feature set is designed around the messy details that make speech systems useful after the first line of text appears.

For developers comparing services, the sensible next step is controlled evaluation: use the same representative clips, score the errors that matter to the product, measure latency at expected volume and check how diarization and terminology handling behave. MAI-Transcribe-2's launch claims are strong. The preview gives teams a way to find out how much of that advantage survives contact with their own audio.

The preview status is another reason to keep conclusions narrow. Microsoft can change pricing, limits or model behavior as the service matures, and application teams may discover edge cases that a benchmark set does not expose. The strongest launch claim is therefore not that every existing speech service has been made obsolete. It is that Microsoft has moved a new internally developed transcription model into a form developers can test against real audio, with the measurements and controls needed to make an evidence-based choice.

Sources

More from Xarmo News