Microsoft releases three MAI foundation models: MAI-Voice-1, MAI-Transcribe-1, MAI-Image-2

TL;DR AI
2 min readKey summary
Microsoft announced three new MAI foundation models—MAI-Transcribe-1, MAI-Voice-1, and MAI-Image-2—available in Microsoft Foundry and MAI Playground.
MAI-Transcribe-1 achieved a 3.9% WER on FLEURS across 25 languages and is reported to be fast.
MAI-Voice-1 can generate 60 seconds of audio from 1 second of input and can create custom voices from a few seconds of audio.
MAI-Image-2 ranks in the top three on Arena.ai and delivered at least 2x faster generation on Foundry and Copilot while maintaining quality.


