grok-imagine-2.0-extGrok Imagine 2.0 Ext 是通过 APIMart 异步图像生成 API 提供的文生图模型。单次请求可按 7 种宽高比之一生成 1–12 张图片,并按实际成功交付的图片张数计费。

Image link valid for 72 hours
透明定价,无隐藏费用。按量付费,用多少付多少。
* 实际费用以最终输出为准。
围绕明确生成参数设计的纯文生图工作流
根据文本为常见创意与产品流程生成视觉素材
配置访问权限、提交文生图任务并轮询结果
登录 APIMart 或创建账号,以访问控制面板。
在发送正式生成请求前,为账户余额充值。
生成 API Key,使用提示词提交 grok-imagine-2.0-ext 请求,再通过返回的任务 ID 轮询图片 URL。
了解输入、生成张数、宽高比、异步任务、质量、URL 与计费规则
它只支持文本生成图像。请提交非空文本提示词;不支持图生图,也不接受图片 URL 或带角色的图片输入。
n 参数接受 1 到 12 的任意整数,默认值为 1。简洁的 UI 可以提供推荐选项 1、4、8、12。
推荐使用 1:1、2:3、3:2、3:4、4:3、9:16 和 16:9 七种比例。不要发送 1:2、2:1、4:5 或 auto 等不支持的值。
向 /v1/images/generations 发送 POST 请求,保存返回的任务 ID,再轮询 /v1/tasks/{task_id}。任务完成后,从结果中读取成功返回的图片 URL。
质量模式固定:使用 resolution=quality 或省略该字段,不能选择 1K、2K 或 4K 档位。结果只提供 URL,且链接会在 72 小时后过期,请及时保存。
按实际成功交付的图片张数计费。如果成功张数少于请求张数,只收取成功图片的费用;失败任务不收费。
探索同类型的其他模型。

Nano banana
Nano Banana (gemini-2.5-flash-image-preview) is Google DeepMind's fast, conversational image generation and editing model. It delivers natural-language text-to-image, precise multi-turn edits, strong character consistency, and multi-image fusion at low latency.

Doubao Seedream 5.0 Pro
Seedream 5.0 Pro (doubao-seedream-5-0-pro) is ByteDance's quality-first text-to-image model. It produces cinematic 1K and 2K images with best-in-class text rendering, supports up to 10 reference images for style and character consistency, and unifies generation and editing in a single API call.

Midjourney
Midjourney is a leading text-to-image AI model renowned for its painterly aesthetics, strong composition, and cinematic lighting. Generate high-quality images from text prompts or reference images, with fine-grained control over style, aspect ratio, and creative parameters — accessible programmatically through APIMart's unified API, no Discord required.

Wan 2.7 Image
wan‑2‑7‑image is an advanced image generation model in Alibaba’s Wan 2.7 series designed to create high‑quality visuals from text prompts. It excels at generating detailed, realistic images with accurate object representation and rich composition. The model supports multimodal input (e.g., combining text with reference images) to influence style and structure, and it’s well‑suited for creative workflows such as marketing assets, product visuals, social media graphics, and artistic content production. With strong semantic understanding and prompt adherence, wan‑2‑7‑image delivers both fidelity and expressive visual output.