Microsoft AI has introduced MAI-Transcribe-2-Streaming, described as its first real-time speech-to-text model. The system is built for streaming transcription, meaning it processes audio as it arrives rather than waiting for a complete recording.

On the Artificial Analysis AA-WER Streaming benchmark, the model ranks first among 38 competing models. It achieves a word error rate of 2.5% at a latency of 0.13 seconds for final transcripts, and the same 2.5% error rate at 0.12 seconds for first partials, according to the source.

The model also supports 60 languages, making it a broadly multilingual option for live transcription. The source does not compare it directly against specific rivals beyond the benchmark ranking, so no further performance claims are available.