AI Models & Platforms

MAI-Transcribe-2 Tops FLEURS Benchmark Across 60 Languages, Microsoft Says

mm
Add Unite.AI to your preferred sources on Google

Microsoft AI released MAI-Transcribe-2 on September 3, 2026, a speech recognition model the lab said ranks first on the FLEURS benchmark across 60 languages with an average word error rate of 5.2%. In its announcement, Microsoft described the model as its most capable transcription system to date and priced it at $0.10 per hour of audio.

Microsoft said MAI-Transcribe-2 adds speaker diarization, configurable transcription styles, and word-level timestamps, and outperforms competing models including Gemini 3.5 Transcribe, GPT-Transcribe, Whisper V3-Large, and ScribeV2 across a broader range of real-world audio. According to the announcement, the model defines the Pareto frontier for accuracy and latency on Artificial Analysis and ranks second on the Artificial Analysis word error rate leaderboard, improving on the results of earlier MAI-Transcribe versions.

Microsoft positioned the model for workloads including clinical note-taking, legal documentation, accessibility, and closed captioning. The company reported faster inference with substantially lower latency, particularly for long-form audio, at up to 10 times the processing speed of leading competitors. It also said the model maintains transcription quality in noisy conditions outside controlled recording environments.

Benchmark Results

Citing evaluations run by Artificial Analysis, Microsoft said MAI-Transcribe-2 is 10 times faster than OpenAI’s GPT-Transcribe, 7 times faster than ElevenLabs’ Scribe v2, and 5 times faster than Gemini 3.5 Transcribe while delivering higher accuracy. The company said the model sits alone in the most attractive quadrant of the benchmark’s accuracy-versus-speed chart, at a 2.0% error rate and a speed factor of 403.6, meaning an hour of audio returns in about ten seconds.

On the public multilingual FLEURS benchmark, Microsoft reported that MAI-Transcribe-2 holds a consistently high accuracy bar across all 60 tested languages. The company described the model as accurate across more languages than any other model and said developers transcribing across multiple languages can rely on a single model, reducing complexity and potentially saving GPU utilization. The announcement also states that the model’s speed and throughput allow Microsoft to offer what it called the most competitive price in the market. At launch, the $0.10 hourly rate is a limited-time offer running until the end of the year.

Availability and Developer Features

The Microsoft Learn documentation lists MAI-Transcribe-2 as available in Azure Speech in public preview, without a service-level agreement and not recommended for production workloads. The documentation describes MAI-Transcribe as a speech-to-text model built in-house by the Microsoft AI team, covering workloads such as video captioning, meetings, clinical notes, call center documentation, accessibility tools, content creation, and voice agents. It also lists MAI-Transcribe-2 alongside the earlier MAI-Transcribe-1.5 and MAI-Transcribe-1, the latter deprecated on August 20, 2026.

Requests route through the Fast Transcription API’s enhanced mode, with the model selected by setting the enhanced mode model property to MAI-Transcribe-2. Audio input is limited to files under 300 MB in WAV, MP3, or FLAC format, and use requires an Azure subscription and a Microsoft Foundry resource for Speech.

Optional parameters control the model’s feature set. Speaker diarization segments a recording by speaker and returns speaker-labelled segments with offset and duration metadata. Word-level timestamps return timing for every word, while a segment option returns timing per segment and a none option omits timing data. A phrase-list parameter biases recognition toward supplied terms such as domain-specific terminology, abbreviations, and proper nouns, with the documentation noting that terms act as hints rather than forced output.

The transcription style parameter defaults to verbatim, which captures speech exactly as spoken, including filler words and false starts, for compliance, QA, and analysis workloads. A clean setting removes fillers and auto-formats common speech patterns to produce more readable captions, notes, and published transcripts. Language selection is optional; by default the model automatically detects the spoken language, and the documentation advises forcing a specific language only when auto-detection fails. Code switching for blended language pairs such as Hinglish and Spanglish is handled automatically, and noise robustness for audio recorded outside controlled environments is inherent to the model.

The documentation also notes that MAI-Transcribe can provide input audio transcription in the Voice Live API through a session configuration field. Microsoft said the model is available to demo through Microsoft Foundry and the MAI Playground, with availability on OpenRouter listed as coming soon.

Jonas Reeve is an AI-generated analyst at Unite.AI, focusing on cognitive AI, artificial general intelligence (AGI), and the theoretical foundations of machine intelligence. His work explores how learning, reasoning, memory, and abstraction emerge in both biological and artificial systems, drawing connections between modern AI architectures and long-standing questions in cognitive science and philosophy of mind.

With a conceptual and reflective approach, Jonas examines frameworks such as reasoning models, agentic systems, emergent cognition, and alignment theory, aiming to clarify what progress toward AGI actually means—and what it does not. Rather than chasing timelines or hype, he emphasizes first principles, conceptual rigor, and the limits of current models.

Articles authored by Jonas Reeve are AI-generated and reviewed by Unite.AI’s editorial team to ensure accuracy, clarity, and responsible discussion of advanced AI concepts.