MAI-Image-2.6は文字描画、ポートレート、3D画像、商品・ブランドビジュアル、フォトリアルな出力を強化します。APIMartがAPIアクセスを準備する間、確認済みの機能をご覧ください。
Microsoft AIが発表した画像生成機能を先行紹介
APIMartはMAI-Image-2.6 APIの統合を準備中です。公開時に、対応入力モード、パラメーター、解像度、上限、料金、アクセス方法を掲載します。
Microsoft AIの発表とArenaデビューで確認された機能
発表済み機能に基づく制作・クリエイティブ用途
確認済みのモデル情報とAPIMartでの提供予定
Microsoft AIの画像生成モデルへプログラムからアクセスするAPIです。文字、ポートレート、3D、商用デザイン、フォトリアル表現の向上が確認されています。
公開時に1,336点で2位でした。MicrosoftはMAI-Image-2.5比で全体79 Elo、文字描画91 Eloの向上を発表しています。
はい。文字描画の強化は公式に確認され、Arenaの同部門で91 Elo向上したと発表されています。
Microsoftは複数参照と豊かなグラウンディングを発表しています。枚数、ファイル要件、APIMartの方式は統合テスト後に公開します。
形式と解像度の制御強化は発表済みですが、完全なパラメーターは未公開です。APIMartは実際に利用できる選択肢のみ記載します。
APIは近日公開予定ですが、正確な日付は未発表です。統合完了時にドキュメント、上限、料金を公開します。
最終的なAPI範囲は未確定です。入力、参照モード、解像度、パラメーター、商用条件、料金は公開時に掲載します。
ページ右下のライブチャット、[email protected] へのメール、またはDiscordコミュニティからお問い合わせいただけます。チームができるだけ早くご対応いたします。
同じカテゴリの他のモデルを探す。

Nano banana
Nano Banana (gemini-2.5-flash-image-preview) is Google DeepMind's fast, conversational image generation and editing model. It delivers natural-language text-to-image, precise multi-turn edits, strong character consistency, and multi-image fusion at low latency.

Doubao Seedream 5.0 Pro
Seedream 5.0 Pro (doubao-seedream-5-0-pro) is ByteDance's quality-first text-to-image model. It produces cinematic 1K and 2K images with best-in-class text rendering, supports up to 10 reference images for style and character consistency, and unifies generation and editing in a single API call.

Midjourney
Midjourney is a leading text-to-image AI model renowned for its painterly aesthetics, strong composition, and cinematic lighting. Generate high-quality images from text prompts or reference images, with fine-grained control over style, aspect ratio, and creative parameters — accessible programmatically through APIMart's unified API, no Discord required.

Wan 2.7 Image
wan‑2‑7‑image is an advanced image generation model in Alibaba’s Wan 2.7 series designed to create high‑quality visuals from text prompts. It excels at generating detailed, realistic images with accurate object representation and rich composition. The model supports multimodal input (e.g., combining text with reference images) to influence style and structure, and it’s well‑suited for creative workflows such as marketing assets, product visuals, social media graphics, and artistic content production. With strong semantic understanding and prompt adherence, wan‑2‑7‑image delivers both fidelity and expressive visual output.