Microsoft AI 的 MAI-Image-2.6 提升了文字渲染、人物肖像、3D 图像、商品与品牌视觉以及照片级输出。APIMart 正在准备 API 接入详情,欢迎先行了解已经确认的能力。
抢先了解 Microsoft AI 已公布的图像生成能力
APIMart 正在准备接入 MAI-Image-2.6 API。上线后,本页将公布支持的输入模式、请求参数、分辨率、限制、价格与访问详情。
Microsoft AI 官方公告与 Arena 首发中已经确认的能力
基于微软已公布能力的创意与制作场景
已经确认的 MAI-Image-2.6 模型信息,以及 APIMart 即将提供的访问能力
MAI-Image-2.6 API 将提供对 Microsoft AI 图像生成模型的程序化访问。微软已公布其在文字渲染、人物肖像、3D、商业设计与照片级图像方面的提升。
MAI-Image-2.6 首发以 1,336 分位列 Arena 文生图榜第 2。微软表示,其整体相比 MAI-Image-2.5 提升 79 Elo,文字渲染类别提升 91 Elo。
可以。更强的文字渲染是官方确认的模型改进;微软表示,相比 MAI-Image-2.5,其 Arena 文字渲染类别提升了 91 Elo。
微软表示,该模型能跨多个参考素材工作,并提供更丰富的语义依据。APIMart 支持的数量、文件要求与具体参考工作流将在接入测试后公布。
微软已宣布增强对格式和分辨率的控制,但尚未公布完整参数。APIMart 将只记录最终集成实际提供的选项。
MAI-Image-2.6 API 即将登陆 APIMart,目前尚未公布准确日期。集成就绪后将发布文档、限制与价格。
MAI-Image-2.6 API 的最终接入范围尚未确认。支持的输入、参考模式、分辨率、参数、商业条款与价格将在上线时公布。
你可以通过页面右下角的在线客服、发送邮件至 [email protected],或在 Discord 社区中联系我们。我们的团队会尽快回复你的问题。
探索同类型的其他模型。

Nano banana
Nano Banana (gemini-2.5-flash-image-preview) is Google DeepMind's fast, conversational image generation and editing model. It delivers natural-language text-to-image, precise multi-turn edits, strong character consistency, and multi-image fusion at low latency.

Doubao Seedream 5.0 Pro
Seedream 5.0 Pro (doubao-seedream-5-0-pro) is ByteDance's quality-first text-to-image model. It produces cinematic 1K and 2K images with best-in-class text rendering, supports up to 10 reference images for style and character consistency, and unifies generation and editing in a single API call.

Midjourney
Midjourney is a leading text-to-image AI model renowned for its painterly aesthetics, strong composition, and cinematic lighting. Generate high-quality images from text prompts or reference images, with fine-grained control over style, aspect ratio, and creative parameters — accessible programmatically through APIMart's unified API, no Discord required.

Wan 2.7 Image
wan‑2‑7‑image is an advanced image generation model in Alibaba’s Wan 2.7 series designed to create high‑quality visuals from text prompts. It excels at generating detailed, realistic images with accurate object representation and rich composition. The model supports multimodal input (e.g., combining text with reference images) to influence style and structure, and it’s well‑suited for creative workflows such as marketing assets, product visuals, social media graphics, and artistic content production. With strong semantic understanding and prompt adherence, wan‑2‑7‑image delivers both fidelity and expressive visual output.