< Back to all clusters
[TECHNOLOGY] · United States, China, India · 8 sources

started · updated

Microsoft launches MAI-Transcribe-2 AI speech model

Microsoft has launched MAI-Transcribe-2, a new AI speech-recognition model designed to convert audio to text with high speed and accuracy. The model supports 60 languages and features advanced capabilities such as speaker diarization, which distinguishes between different voices in a single recording, and word-level timestamps.

In terms of performance, Microsoft claims the model is up to 10 times faster than OpenAI’s GPT-Transcribe and five times faster than Google’s Gemini 3.5 Transcribe. It achieved an average Word Error Rate (WER) of 5.2% on the FLEURS benchmark.

The service is positioned as a highly cost-effective solution for enterprises, with introductory pricing set at approximately $0.10 per hour of audio. This represents a significant reduction from previous models, aiming to provide substantial savings for businesses processing large volumes of audio data, such as call centers.

Entities

Google · MAI-Transcribe-2 · Microsoft · OpenAI