< Back to situations

Monitor this situation.

[SITUATION] · [QUIET] · [TECHNOLOGY]

2 clusters · 10 sources · 1 days · First seen · Last updated

Microsoft AI model releases

Overview

Microsoft has expanded its AI capabilities through the release of several new models. The company first launched MAI-Transcribe-2, an AI transcription model designed for high-speed, multilingual speech-to-text tasks. Microsoft claims the model is up to 10 times faster than OpenAI’s GPT-Transcribe and five times faster than Google’s Gemini 3.5 Transcribe, featuring a 5.2% average Word Error Rate across 60 languages.

MAI-Transcribe-2 includes advanced capabilities such as speaker diarization to distinguish between different voices and word-level timestamps. Positioned as a cost-effective enterprise solution, the service has introductory pricing of approximately $0.10 per hour of audio, aimed at providing savings for high-volume users like call centers.

Following this, Microsoft released the MAI-Image-2.6 and MAI-Image-2.6-Flash models via its Foundry developer platform. These models support text-to-image generation and editing, with the Flash version optimized for high-volume, high-speed applications. While the models introduce new pricing structures for text and image tokens, reports have noted discrepancies in Microsoft’s official documentation regarding performance metrics and rankings in industry comparisons.

Entities

OpenAI · Microsoft · Google · Microsoft AI · MAI-Transcribe-2

Timeline

  1. 7 days ago

    [TECHNOLOGY] 2 sources
    Microsoft launches MAI-Image-2.6 AI image models in Foundry

    Microsoft has launched MAI-Image-2.6 and MAI-Image-2.6-Flash on its Foundry platform, though documentation shows conflicting performance and ranking data.

  2. 8 days ago

    [TECHNOLOGY] 8 sources
    Microsoft launches MAI-Transcribe-2 AI speech model

    Microsoft launched MAI-Transcribe-2, an AI speech-to-text model that is faster, more accurate, and cheaper than competitors from OpenAI and Google, supporting 60 languages and advanced speaker identification.

Sources

4sysops.com · antaranews.com · blogspan.net · industry.co.id · ithome.com · mediaindonesia.com · newsbytesapp.com · pemilu2024.harianjogja.com · pendidikan.harianjogja.com · womanindonesia.co.id

This summary has been updated 1 time: see revision history