Microsoft Builds Its Own AI Model Stack To Reduce OpenAI Dependence

Key summary
Microsoft released three in-house models—MAI-Transcribe-1, MAI-Voice-1 and MAI-Image-2—for commercial use via Foundry and the MAI Playground.
The models cover speech transcription, voice generation and image creation.
Microsoft says MAI-Transcribe-1 achieves the lowest average word error rate on the FLEURS benchmark for the top 25 languages by Microsoft product usage, outperforms OpenAI's Whisper-large-v3 across those languages, beats Google's Gemini 3.1 Flash on a majority of the remaining languages, and offers faster batch transcription.
The move is meant to hedge Microsoft’s dependence on OpenAI; an October 2025 restructuring gave Microsoft the right to pursue AGI independently or with other partners, reduced its OpenAI stake from 32.5% to about 27%, extended Microsoft’s IP rights through 2032, removed Microsoft’s exclusive compute-provider role, and left OpenAI able to work with other clouds while committing an additional $250 billion in Azure purchases.
Mustafa Suleyman leads the Microsoft AI Superintelligence team full-time as of November 2025; the first MAI models shipped in August 2025 and MAI-Image-1 arrived in October 2025.


