새로운 시각 콘셉트를 만들고 선택 영역을 수정하며, 반복되는 피사체의 일관성과 외부 시각 맥락을 다단계 제작 과정에서 활용할 수 있습니다.
시각 생성과 이미지 편집, 검색 기반 콘셉트 및 반복 디자인 작업을 위한 멀티모달 모델을 살펴봅니다.
APIMart는 Gemini Nano Banana 2.1 API 연동을 준비하고 있습니다. 연동 준비가 완료되면 이 페이지에서 지원 모델 ID, 요청 형식, 제한, 요금 및 이용 방법을 안내합니다.
생성, 이미지 편집, 검색 연동, 일관성과 실용적인 제작 제어
크리에이티브, 커머스, 제품과 콘텐츠 팀을 위한 이미지 편집 워크플로
핵심 기능, 입력, 편집 흐름, 제어와 APIMart 접근 정보
이미지 생성, 이미지 편집, 검색 기반 시각 제작과 반복 수정을 위해 설계된 Google의 멀티모달 이미지 모델입니다.
텍스트와 이미지를 입력할 수 있고 동영상도 입력 맥락으로 활용하므로 지시와 시각 참조를 결합한 이미지 작업을 구성할 수 있습니다.
여러 대화 턴에 걸쳐 이미지 편집을 이어 갈 수 있으며 마스크 기반 수정으로 기존 시각물의 선택 영역에 변경을 집중할 수 있습니다.
피사체 일관성이 주요 초점이므로 관련 이미지에서 같은 사람, 제품 또는 사물을 반복해서 다루는 작업에 적합합니다.
Google 검색과 이미지 검색 그라운딩은 최신 정보나 현실의 시각 자료가 필요한 이미지 작업에 외부 맥락을 제공합니다.
문서에는 1K, 2K, 4K 출력과 정사각형, 세로형, 가로형, 파노라마를 포함한 다양한 화면 비율이 안내되어 있습니다.
temperature는 지원되지 않습니다. topP와 topK뿐 아니라 seed와 logprobs도 사용할 수 없으므로 문서화된 요청 형식을 따라야 합니다.
APIMart는 Gemini Nano Banana 2.1 API 연동을 준비하고 있습니다. 준비가 완료되면 지원 모델 ID, 요청 형식, 제한, 요금 및 이용 방법을 이 페이지에서 안내합니다.
같은 카테고리의 다른 모델을 탐색하세요.

Nano banana
Nano Banana (gemini-2.5-flash-image-preview) is Google DeepMind's fast, conversational image generation and editing model. It delivers natural-language text-to-image, precise multi-turn edits, strong character consistency, and multi-image fusion at low latency.

Seedream 5.0 Pro
Seedream 5.0 Pro (seedream-5-0-pro) is quality-first text-to-image model. It produces cinematic 1K and 2K images with best-in-class text rendering, supports up to 10 reference images for style and character consistency, and unifies generation and editing in a single API call.

Midjourney
Midjourney is a leading text-to-image AI model renowned for its painterly aesthetics, strong composition, and cinematic lighting. Generate high-quality images from text prompts or reference images, with fine-grained control over style, aspect ratio, and creative parameters — accessible programmatically through APIMart's unified API, no Discord required.

Wan 2.7 Image
wan‑2‑7‑image is an advanced image generation model in Alibaba’s Wan 2.7 series designed to create high‑quality visuals from text prompts. It excels at generating detailed, realistic images with accurate object representation and rich composition. The model supports multimodal input (e.g., combining text with reference images) to influence style and structure, and it’s well‑suited for creative workflows such as marketing assets, product visuals, social media graphics, and artistic content production. With strong semantic understanding and prompt adherence, wan‑2‑7‑image delivers both fidelity and expressive visual output.