生成新的视觉概念、进行指定区域的图像 编辑、维持重复主体的一致性,并在多阶段创作中利用外部视觉信息和参考素材。
了解一款适合视觉生成、精细图像 编辑、搜索溯源概念和迭代设计工作的多模态模型。
APIMart 正在准备接入 Gemini Nano Banana 2.1 API。集成就绪后,本页将公布支持的模型 ID、请求格式、限制、价格与访问详情。
覆盖生成、图像 编辑、搜索溯源、一致性和实用创作控制
适合创意、电商、产品与内容团队的实用图像 编辑工作流
了解核心能力、输入、编辑流程、控制方式及APIMart接入信息
它是Google推出的多模态图像模型,面向图像生成、图像编辑、搜索溯源视觉创作及连续迭代。
模型可使用文本和图像输入,也能将视频作为输入上下文,从而把文字指令与视觉参考组合到同一图像任务中。
用户可在多个对话轮次中继续图像 编辑,并通过蒙版修改把变化集中到现有视觉内容的指定区域。
主体一致性是其重点方向,适合在相关图像中反复呈现同一人物、商品或物体的创作流程。
Google搜索和图片搜索可为需要近期信息或现实视觉资料的图像任务提供外部上下文。
公开文档列出1K、2K和4K输出,并提供覆盖方形、纵向、横向及全景布局的多种宽高比。
模型不支持temperature。topP与topK也不可用,seed和logprobs同样不能发送;应用应遵循公开的请求格式。
APIMart正在准备接入Gemini Nano Banana 2.1 API。集成就绪后,本页将公布支持的模型ID、请求格式、限制、价格与访问详情。
探索同类型的其他模型。

Nano banana
Nano Banana (gemini-2.5-flash-image-preview) is Google DeepMind's fast, conversational image generation and editing model. It delivers natural-language text-to-image, precise multi-turn edits, strong character consistency, and multi-image fusion at low latency.

Seedream 5.0 Pro
Seedream 5.0 Pro (seedream-5-0-pro) is quality-first text-to-image model. It produces cinematic 1K and 2K images with best-in-class text rendering, supports up to 10 reference images for style and character consistency, and unifies generation and editing in a single API call.

Midjourney
Midjourney is a leading text-to-image AI model renowned for its painterly aesthetics, strong composition, and cinematic lighting. Generate high-quality images from text prompts or reference images, with fine-grained control over style, aspect ratio, and creative parameters — accessible programmatically through APIMart's unified API, no Discord required.

Wan 2.7 Image
wan‑2‑7‑image is an advanced image generation model in Alibaba’s Wan 2.7 series designed to create high‑quality visuals from text prompts. It excels at generating detailed, realistic images with accurate object representation and rich composition. The model supports multimodal input (e.g., combining text with reference images) to influence style and structure, and it’s well‑suited for creative workflows such as marketing assets, product visuals, social media graphics, and artistic content production. With strong semantic understanding and prompt adherence, wan‑2‑7‑image delivers both fidelity and expressive visual output.