Microsoft Launches Its Strongest AI Image Generation Model MAI-Image-2.5
Published · May 28 · Thu Source · IT之家 (CN)

Microsoft Launches Its Strongest AI Image Generation Model MAI-Image-2.5

Microsoft Research has launched its strongest image generation model, MAI-Image-2.5, which has risen to 3rd place on the Arena text-to-image leaderboard. This breaks the previous situation where the top 5 were occupied solely by Google DeepMind and OpenAI, representing a 72-point improvement over the previous generation. The model focuses on enhancing text rendering capabilities, making it suitable for generating commercial materials such as infographics and posters, while also showing significant progress in visual reasoning, image details, and style consistency.

KeywordsOpenAIGoogleMicrosoftLaunchesItsStrongestAIImage

Last year, Microsoft announced the launch of its self-developed image generation model, MAI-Image-Annual-1. When the model first appeared in Arena's Image Arena, it ranked only 9th, significantly lagging behind models from other AI laboratories. Subsequently, Microsoft made the model available to users of the Bing.com/create website and the Bing mobile app.

In March this year, the Microsoft AI team launched the second-generation image generation model, MAI-Image-2. The model achieved significant improvements, capable of generating images with more realistic natural lighting, more accurate skin tones, and other effects. MAI-Image-2 achieved a 3rd place ranking upon its debut, second only to Google's gemini-3.1-flash-image-preview and OpenAI's gpt-image-1.5-high-fidelity.

The model is integrated into Copilot and Bing Image Creator, and is also open to developers via the Microsoft Foundry API.

MAI-Image-2.5

Today, the Microsoft AI team announced the launch of its latest text-to-image model, MAI-Image-2.5. According to Arena's latest text-to-image leaderboard, the model currently ranks third. Currently, the top spot is held by OpenAI's gpt-image-2, with a score of 1388.

According to Microsoft, the new MAI-Image-2.5 model performs better across a wide range of image styles. The model is designed to follow prompts more closely, render text more reliably, and generate images with richer details and stronger coherence. Microsoft also stated that the model possesses stronger visual reasoning capabilities, allowing it to better understand objects, lighting, scale, scene structure, and spatial relationships.

Microsoft specifically emphasized that the new model has achieved the greatest improvements in text rendering, stylized illustrations, and commercial images. This enables users to generate higher-quality posters, packaging renders, brand concept images, and product promotional images. Text in generated images will be clearer and sharper, layouts will be more stable, and brand-centric visual elements will be presented more refinedly.

Like previous Microsoft AI models, the new MAI-Image-2.5 model is open for anyone to try on Arena today. It is expected that the new image generation model will also go live on MAI Playground and Microsoft Foundry within the next two weeks. Return to Sohu, view more.

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.