Microsoft's New AI Models Go Beyond Just Text

TL;DR AI
2 min readKey summary
Microsoft announced three new AI models: a voice model, a text transcription model, and a second-generation image model.
The transcription model converts recordings into text in 25 languages and targets captioning, meetings, and voice agents.
The voice model can produce up to 60-second audio clips and MAI-Image-2 offers faster, more lifelike image generation.
The new models are available in Foundry and the MAI playground, with plans to add MAI-Image-2 to Bing and PowerPoint.


