MAI-Image-2.6은 텍스트, 인물, 3D 이미지, 제품·브랜드 비주얼과 포토리얼 출력을 개선합니다. APIMart가 API 액세스를 준비하는 동안 확인된 기능을 살펴보세요.
Microsoft AI가 발표한 이미지 생성 기능을 미리 확인하세요
APIMart는 MAI-Image-2.6 API 통합을 준비하고 있습니다. 출시 시 지원 입력 모드, 파라미터, 해상도, 제한, 가격과 액세스 정보를 공개합니다.
Microsoft AI 발표와 Arena 데뷔에서 확인된 기능
발표된 기능을 기반으로 한 크리에이티브 및 제작 활용 사례
확인된 모델 정보와 예정된 APIMart 액세스
Microsoft AI 이미지 생성 모델에 프로그래밍 방식으로 접근하는 API입니다. 텍스트, 인물, 3D, 상업 디자인과 포토리얼 이미지 향상이 확인됐습니다.
출시 시 1,336점으로 2위였습니다. Microsoft는 MAI-Image-2.5 대비 전체 79 Elo, 텍스트 부문 91 Elo 향상을 발표했습니다.
네. 텍스트 렌더링 향상은 공식 확인됐으며 Arena 해당 부문에서 91 Elo 개선됐습니다.
Microsoft는 여러 참조와 풍부한 그라운딩을 발표했습니다. 수량, 파일 요건과 APIMart 워크플로는 통합 테스트 후 공개됩니다.
형식과 해상도 제어 향상은 발표됐지만 전체 파라미터는 미공개입니다. APIMart는 실제 제공 옵션만 문서화합니다.
API는 곧 출시될 예정이지만 정확한 날짜는 발표되지 않았습니다. 통합 완료 시 문서, 제한과 가격을 공개합니다.
최종 API 범위는 미확정입니다. 입력, 참조 모드, 해상도, 파라미터, 상업 조건과 가격은 출시 시 공개됩니다.
페이지 오른쪽 하단의 실시간 채팅, [email protected] 이메일, 또는 Discord 커뮤니티를 통해 문의하실 수 있습니다. 최대한 빠르게 답변드리겠습니다.
같은 카테고리의 다른 모델을 탐색하세요.

Nano banana
Nano Banana (gemini-2.5-flash-image-preview) is Google DeepMind's fast, conversational image generation and editing model. It delivers natural-language text-to-image, precise multi-turn edits, strong character consistency, and multi-image fusion at low latency.

Doubao Seedream 5.0 Pro
Seedream 5.0 Pro (doubao-seedream-5-0-pro) is ByteDance's quality-first text-to-image model. It produces cinematic 1K and 2K images with best-in-class text rendering, supports up to 10 reference images for style and character consistency, and unifies generation and editing in a single API call.

Midjourney
Midjourney is a leading text-to-image AI model renowned for its painterly aesthetics, strong composition, and cinematic lighting. Generate high-quality images from text prompts or reference images, with fine-grained control over style, aspect ratio, and creative parameters — accessible programmatically through APIMart's unified API, no Discord required.

Wan 2.7 Image
wan‑2‑7‑image is an advanced image generation model in Alibaba’s Wan 2.7 series designed to create high‑quality visuals from text prompts. It excels at generating detailed, realistic images with accurate object representation and rich composition. The model supports multimodal input (e.g., combining text with reference images) to influence style and structure, and it’s well‑suited for creative workflows such as marketing assets, product visuals, social media graphics, and artistic content production. With strong semantic understanding and prompt adherence, wan‑2‑7‑image delivers both fidelity and expressive visual output.