Microsoft AI's MAI-Image-2.6 improves text rendering, portraits, 3D imagery, product and branding visuals, and photorealistic output. Explore the confirmed capabilities while APIMart prepares API access details.
A first look at the image-generation capabilities announced by Microsoft AI
APIMart is preparing the MAI-Image-2.6 API integration. At launch, this page will publish supported input modes, request parameters, resolutions, limits, pricing, and access details.
Confirmed capabilities from Microsoft AI's announcement and the Arena debut
Creative and production scenarios supported by the capabilities Microsoft has announced
Confirmed MAI-Image-2.6 model information and upcoming APIMart access
The MAI-Image-2.6 API will provide programmatic access to Microsoft AI's image-generation model. Microsoft announced improvements in text rendering, portraits, 3D, commercial design, and photorealistic imagery.
At launch, MAI-Image-2.6 ranked No. 2 on Arena's Text-to-Image leaderboard with 1,336 points. Microsoft reports a 79-Elo overall gain over MAI-Image-2.5 and a 91-Elo gain in text rendering.
Yes. Stronger text rendering is a confirmed model improvement, with Microsoft reporting a 91-Elo increase in Arena's text-rendering category compared with MAI-Image-2.5.
Microsoft says the model can work across multiple references with richer grounding. The exact number, file requirements, and supported reference workflows for APIMart will be published after integration testing.
Microsoft has announced greater control over format and resolution but has not published the full parameter set. APIMart will document only the options available through its final integration.
MAI-Image-2.6 API access is coming soon to APIMart, but an exact date has not been announced. Documentation, limits, and pricing will be published when the integration is ready.
The final MAI-Image-2.6 API scope has not been confirmed. Supported inputs, reference modes, resolutions, parameters, commercial terms, and pricing will be published at launch.
You can reach us via the live chat in the bottom-right corner, email us at [email protected], or join our Discord community. Our team will get back to you as soon as possible.
Explore more models in the same category.

Nano banana
Nano Banana (gemini-2.5-flash-image-preview) is Google DeepMind's fast, conversational image generation and editing model. It delivers natural-language text-to-image, precise multi-turn edits, strong character consistency, and multi-image fusion at low latency.

Doubao Seedream 5.0 Pro
Seedream 5.0 Pro (doubao-seedream-5-0-pro) is ByteDance's quality-first text-to-image model. It produces cinematic 1K and 2K images with best-in-class text rendering, supports up to 10 reference images for style and character consistency, and unifies generation and editing in a single API call.

Midjourney
Midjourney is a leading text-to-image AI model renowned for its painterly aesthetics, strong composition, and cinematic lighting. Generate high-quality images from text prompts or reference images, with fine-grained control over style, aspect ratio, and creative parameters — accessible programmatically through APIMart's unified API, no Discord required.

Wan 2.7 Image
wan‑2‑7‑image is an advanced image generation model in Alibaba’s Wan 2.7 series designed to create high‑quality visuals from text prompts. It excels at generating detailed, realistic images with accurate object representation and rich composition. The model supports multimodal input (e.g., combining text with reference images) to influence style and structure, and it’s well‑suited for creative workflows such as marketing assets, product visuals, social media graphics, and artistic content production. With strong semantic understanding and prompt adherence, wan‑2‑7‑image delivers both fidelity and expressive visual output.