Physix Frontier · News Briefing Card (IT Home · Oct 2, 2026)

Microsoft Debuts First Streaming Transcription Model: 2.5% WER, 0.13s Latency

KEY FACTS

  • Microsoft launched MAI-Transcribe-2-Streaming on October 1, its first real-time streaming speech transcription model.
  • The model supports 60 languages and features automatic continuous language detection.
  • Promotional pricing for the model is $0.54 per hour, roughly $9 per thousand minutes.
  • In Artificial Analysis testing, it achieved a final word error rate of 2.50% with 0.13 seconds of latency.
  • That result ranks first among 28 streaming speech-to-text models.

KEY DATA

2.50%Word error rate
0.13 secondsFinal transcription latency
60Supported languages
$0.54/hourPromotional price

PHYSIX OBSERVATION

Streaming transcription pushes latency down to the hundred-millisecond level, meaning voice interaction shifts from "speak, then understand" to "understand while speaking." Microsoft is courting developers with low pricing plus a top-of-the-leaderboard score, and real-time captions and customer-service agents stand to benefit first. But non-streaming pricing is only $0.10, so users need to weigh cost against their use case.

Source: IT Home report