Microsoft’s ASR Model Posts 2.4% WER and Runs 5x Faster

If your transcription pipeline is burning compute time on long audio files, you are already paying a tax that the competition stopped paying last week.

Manual transcription review across 43 languages was the bottleneck nobody wanted to budget for

Teams handling multilingual audio at scale face two compounding problems: models that are accurate but slow, and models that are fast but fall apart on accents, dialects, and noise. Both failures land in the same place: a human editor cleaning up text that should have been clean already.

One model takes the audio in, returns clean text across 43 languages

Microsoft AI Introduces MAI accepts audio input and returns transcribed text through Microsoft Foundry or natively inside Copilot, Teams, GitHub, and Dynamics 365 Contact Centre. You route audio through the API or the integrated product, and the model returns text with no language pre-selection required. The single-system architecture is the detail that matters most: one model handles 43 languages, including 10 South Asian languages added in this release, without needing separate pipelines per locale.

The people carrying the most transcription debt feel this first

  • Contact centre operations leads at multinational firms who pay per-minute for transcription correction on regional-accent calls they can never fully cover
  • Localization engineers managing subtitle pipelines across South Asian or European language sets who currently split workloads across multiple models
  • Legal and compliance teams who treat high WER as a liability risk on recorded depositions or regulatory call logs

The language expansion alone changes the build-vs-buy calculus for anyone who previously stitched together regional models to cover Bengali, Tamil, Ukrainian, or Catalan.

Whisper held this benchmark for two years and now it does not

OpenAI Whisper became the default reference point for production ASR, but MAI-Transcribe-1.5 now claims first place on the FLEURS multilingual benchmark and third place on the Artificial Analysis leaderboard at 2.4% WER, while running up to 5x faster on long-form audio. If speed-times-accuracy is your actual production constraint, the leaderboard just reshuffled.

What you can do with it starting now

  • Transcribe long customer service calls at 5x the speed of previous benchmarks
  • Process multilingual audio in a single API call without language routing logic
  • Drop WER below 3% on English production workloads without model fine-tuning
  • Cover South Asian and European language sets inside Teams meetings natively

Pricing not listed — check our directory.

The 5x speed claim has a ceiling that real-world audio will find

The 5x speed improvement applies to long-audio transcription specifically, and performance on short, noisy, or heavily accented audio in the 18 newly added languages has not yet been independently replicated outside Microsoft’s own benchmarks.

If MAI-Transcribe-1.5 is not the right fit

AssemblyAI offers strong async transcription with speaker diarization built in, which MAI-Transcribe-1.5 does not yet surface as a standalone feature. Deepgram is worth comparing on latency for real-time streaming use cases where batch throughput is less relevant than time-to-first-word.

The ASR leaderboard is moving faster than most teams are updating their vendor contracts

The gap between what production pipelines are running and what the current benchmarks support is growing every quarter. We cover tools like this every Friday — subscribe here and we’ll send the best ones straight to you.