
Image link valid for 72 hours
隠れた費用のない透明な料金体系。使った分だけお支払い。
* 実際の費用は最終出力に基づきます。
明確な生成パラメーターを備えた、テキスト画像生成専用のワークフロー
一般的なクリエイティブやプロダクト業務向けのビジュアルをテキストから生成
アクセスを設定し、テキスト画像生成タスクを送信して結果をポーリング
APIMart にログインするか、アカウントを作成してダッシュボードを利用します。
本番の生成リクエストを送信する前に、アカウント残高をチャージします。
API キーを発行し、テキストプロンプトと grok-imagine-2.0-ext を送信して、返されたタスク ID から画像 URL を取得します。
入力、生成枚数、アスペクト比、タスク、品質、URL、課金について
テキスト画像生成のみに対応しています。空でないテキストプロンプトを送信してください。画像から画像への生成、画像 URL、役割付き画像の入力には対応していません。
n パラメーターには 1~12 の整数を指定でき、初期値は 1 です。わかりやすい UI では、推奨値の 1、4、8、12 を選択肢にできます。
推奨される 7 つの値は 1:1、2:3、3:2、3:4、4:3、9:16、16:9 です。1:2、2:1、4:5、auto などの非対応値は送信しないでください。
/v1/images/generations に POST し、返されたタスク ID を保存して /v1/tasks/{task_id} をポーリングします。完了したら、結果から正常に返された画像 URL を取得します。
品質は固定です。resolution=quality を使うか省略し、1K、2K、4K の段階は選べません。結果は URL のみで、72 時間後に期限切れになるため、期限内に保存してください。
実際に正常配信された画像の枚数に応じて課金されます。成功枚数が依頼枚数より少ない場合は成功分のみが課金され、失敗したタスクには課金されません。
同じカテゴリの他のモデルを探す。

Nano banana
Nano Banana (gemini-2.5-flash-image-preview) is Google DeepMind's fast, conversational image generation and editing model. It delivers natural-language text-to-image, precise multi-turn edits, strong character consistency, and multi-image fusion at low latency.

Doubao Seedream 5.0 Pro
Seedream 5.0 Pro (doubao-seedream-5-0-pro) is ByteDance's quality-first text-to-image model. It produces cinematic 1K and 2K images with best-in-class text rendering, supports up to 10 reference images for style and character consistency, and unifies generation and editing in a single API call.

Midjourney
Midjourney is a leading text-to-image AI model renowned for its painterly aesthetics, strong composition, and cinematic lighting. Generate high-quality images from text prompts or reference images, with fine-grained control over style, aspect ratio, and creative parameters — accessible programmatically through APIMart's unified API, no Discord required.

Wan 2.7 Image
wan‑2‑7‑image is an advanced image generation model in Alibaba’s Wan 2.7 series designed to create high‑quality visuals from text prompts. It excels at generating detailed, realistic images with accurate object representation and rich composition. The model supports multimodal input (e.g., combining text with reference images) to influence style and structure, and it’s well‑suited for creative workflows such as marketing assets, product visuals, social media graphics, and artistic content production. With strong semantic understanding and prompt adherence, wan‑2‑7‑image delivers both fidelity and expressive visual output.