Image to Video API: 20+ AI Video Models
Image to Video API for developers: animate a photo, product shot or keyframe with Kling V3, Veo 3.1, Seedance 2.5, Wan 3.0, MiniMax Hailuo and more through one API key, async tasks and pricing 20% off list price.
Image to Video API Models on APIMart
Every model below accepts an input image as a first frame or reference and shares the same APIMart key, async task workflow and billing. Open a model for image limits, durations, pricing and code examples.
FLUX 3 Video API
FLUX 3 Video by Black Forest Labs supports text-to-video, image-to-video with 1–10 ordered keyframes, video continuation, and draft enhancement. Generate synchronized audio with 5–20 second HD or FHD output, configurable safety tolerance, and results stored on the APIMart CDN.
Native Audio5–20 SecondsOrdered KeyframesHD and FHDView API & pricing →Wan 3.0 API
Wan 3.0 API (Wan API) provides access to Alibaba Cloud's all-in-one video model. Wan API offers general, first-and-last-frame, and file or webpage generation. Wan API general generation accepts prompts with optional image, video, and audio references. Wan API frame generation requires a first frame and allows an optional last frame. One asynchronous Wan API endpoint creates 2–30 second videos at 480P, 720P, or 1080P with optional output audio.
Three Generation EntriesUp to 30 Seconds480P–1080PMultimodal ReferencesView API & pricing →Veo 3.1 API
Veo 3.1 API on APIMart provides access to both veo3.1-fast and veo3.1-quality models, the newest generation of the Google Veo 3 API. Veo 3.1 is Google DeepMind's upgrade to Veo 3, offering enhanced audiovisual quality, native synchronized audio, and advanced cinematic controls for professional video creation.
Commercial UseView API & pricing →Gemini Omni 1.1 Flash API
Gemini Omni 1.1 Flash is Google's GA Gemini Omni API model for multimodal video, supporting generation, natural-language editing, frame interpolation, and video extension. APIMart provides it through a dedicated official channel with asynchronous task tracking and resolution-based billing. Through the Omni Flash API, Gemini Omni Flash is billed by output resolution, and the live table on this page shows current Google Omni pricing for each tier.
Multimodal VideoFirst-Last Frame360p to 4KView API & pricing →Gemini Omni Flash Preview API
Gemini Omni Flash Preview API on APIMart provides streamlined text-to-video, image-to-video, and video-to-video generation with the Gemini-Omni-Flash-Preview model. Send model and prompt through APIMart's unified video generation endpoint, making Gemini-Omni-Flash-Preview easy to test, integrate, and scale.
Text to VideoImage to VideoVideo to VideoView API & pricing →PixVerse V6 API
PixVerse V6 API on APIMart gives developers access to a video model focused on precise control and native artistic expression. Use pixverse v6 for cinematic control, physics-aware motion, high-fidelity portraits, dynamic aesthetics, reference media guidance, and production-ready AI video workflows through one unified endpoint.
Text to VideoImage to VideoPrecise ControlView API & pricing →Kling V3 API
Kling V3 API on APIMart offers Kuaishou's latest AI video generation models — kling-v3-omni's all-in-one capabilities with multi-modal inputs, and kling-v3's cinematic excellence with up to 15-second video generation. Kling AI 3.0 is offered here as the Kling 3.0 API, so one APIMart key covers both models.
Commercial UseHigh QualityMulti-ModelView API & pricing →Kling V2.6 API
Kling 2.6 API on APIMart delivers battle-tested AI video generation with excellent cost-performance ratio. Features include audio generation, dynamic masks, camera control, and negative prompts for precise creative control.
Commercial UseAudio SupportCamera ControlView API & pricing →Kling 3.0 Turbo API
Kling 3.0 Turbo is Kuaishou's speed-optimized Kling AI video model on APIMart, built for an excellent speed-to-quality ratio. Tuned for low-latency production, it supports text-to-video, image-to-video, flexible aspect ratios, and negative prompts.
Commercial UseLow LatencyFast GenerationView API & pricing →Kling Video O1 API
Kling Video O1 API on APIMart is Kuaishou's flagship AI video generation model. Powered by kling-video-o1's thinking-driven generation, it delivers the highest quality output through advanced reasoning capabilities for unmatched visual fidelity.
Flagship ModelThinking-DrivenHighest QualityView API & pricing →MiniMax Hailuo 02 API
MiniMax Hailuo 02 on APIMart is the latest generation AI video model, offering superior visual quality, high motion coherence, and professional-grade performance. Hailuo 02 (sometimes written Hailuo 2) uses the model ID MiniMax-Hailuo-02 in API requests.
Commercial UseHigh QualityView API & pricing →MiniMax Hailuo 2.3 API
MiniMax Hailuo 2.3 on APIMart is the latest generation AI video model, offering superior visual quality, high motion coherence, and professional-grade performance. Hailuo 2.3 is available as two model IDs, MiniMax-Hailuo-2.3 and MiniMax-Hailuo-2.3-Fast, with one API key.
Commercial UseHigh QualityView API & pricing →MiniMax H3 API
MiniMax H3 on APIMart is the latest generation AI video model, offering native 2K output, high motion coherence, multimodal references, and professional-grade performance.
Commercial UseHigh QualityView API & pricing →Vidu Q3 API
The Vidu Q3 Series offers two powerful AI video generation models: Pro delivers cinematic quality and stunning visual fidelity, while Turbo balances speed and quality for rapid iteration and bulk production.
Commercial UseHigh QualityFast GenerationView API & pricing →Vidu Q4 Preview API
Access the Vidu Q4 Preview API on APIMart with model ID viduq4-preview. Vidu supports one first frame or 1–15 reference images with optional reference audio, producing 3–16 second clips at up to 4K. Generated audio and seed control are available; it does not support text-only generation or end-frame control.
Image to VideoReference to Video3–16 SecondsUp to 4KView API & pricing →Wan 2.5 API
WAN 2.5 API on APIMart is a powerful AI video generation model, offering high visual quality, strong motion coherence, and reliable performance. Alibaba Wan 2.5 handles both text-to-video and image-to-video through one API.
Commercial UseHigh QualityView API & pricing →Wan 2.6 API
WAN 2.6 API on APIMart is the latest generation AI video model, offering superior visual quality, high motion coherence, and professional-grade performance. The wan2.6 API on APIMart includes three model IDs—wan2.6, wan2.6-i2v and wan2.6-i2v-flash—all callable with one key.
Commercial UseHigh QualityView API & pricing →Wan 2.7 API
WAN 2.7 API on APIMart is the latest generation AI video model, offering superior visual quality, high motion coherence, and professional-grade performance. For Wan I2V (image-to-video), set a first frame and an optional last frame on wan2.7, or use wan2.7-r2v for reference-driven video.
Commercial UseHigh QualityView API & pricing →Grok Imagine Video API
Grok Imagine Video API on APIMart handles text-to-video and Grok image to video with xAI's video models. Send a prompt, or add reference image URLs to animate a still image, and get high-quality videos with remarkable visual fidelity. This Grok video API exposes three Grok video model IDs—grok-imagine-video, grok-imagine-video-1.5 and grok-imagine-1.5-video-ext—under one APIMart key.
Video GenerationHigh QualityView API & pricing →SkyReels V4 API
The SkyReels V4 API on APIMart is Skywork AI's first unified video + audio foundation model. SkyReels V4 delivers 1080p at 32 FPS with 15-second clips, native audio, and cinematic multi-shot coherence in a single generation pass.
Video + Audio1080p / 32 FPSView API & pricing →Seedance 1.0 Pro API
Experience Seedance 1.0 Pro Fast video generation technology. Generate stunning videos with logical motion and high fidelity from text or images.
Commercial UseView API & pricing →Seedance 1.5 Pro API
Experience the power of Seedance 1.5 Pro video generation technology. Compared to previous Seedance versions, it offers smoother motion, higher image quality, and stronger instruction following capabilities.
Commercial UseNewSeedance 1.5 ProView API & pricing →Seedance 2.0 API
Seedance 2.0 API on APIMart is ByteDance's latest AI video generation model, delivering stunning visual quality, smooth motion, and professional-grade cinematic output. One Seedance API key gives you every Seedance 2 API variant—seedance-2.0, seedance-2.0-fast and seedance-2.0-mini.
Commercial UseHigh QualityView API & pricing →Seedance 2.5 API
Seedance 2.5 API exposes ByteDance's next-generation audio-video foundation model for production AI video with native 30-second single-pass clips, up to 50 multimodal references, first/last-frame control, region-level edits, and stronger continuity for ads, e-commerce, and film previs.
AI Video30s ClipsMultimodal ReferenceNative AudioView API & pricing →HappyHorse 1.0 API
HappyHorse 1.0 is Alibaba ATH-AI's unified multimodal video model, available on APIMart as the HappyHorse API. It generates 1080p video with synchronized native audio in a single forward pass, supports 7-language lip-sync, and ranks #1 on Artificial Analysis text-to-video and image-to-video leaderboards. Also written as Happy Horse 1.0, this HappyHorse AI video model sits next to the newer HappyHorse 1.1 in the same playground.
Text to VideoImage to VideoNative Audio1080pView API & pricing →
How to Choose an Image to Video API
First and last frame control
Wan 3.0 takes a required first frame with an optional last frame, while Seedance 2.5, Gemini Omni 1.1 Flash, Kling V3 and MiniMax Hailuo also accept start and end frames to control how a shot begins and ends.
Multiple reference images
Seedance 2.5 accepts up to 50 multimodal references, FLUX 3 Video uses 1–10 ordered keyframes, Grok Imagine Video takes up to 7 images and Kling Video O1 up to 2 for style and character control.
Subject and character consistency
Kling V3, Seedance and Wan models animate a reference image while keeping the subject recognizable, a good fit for product shots, avatars and e-commerce clips.
Image to video with native audio
Veo 3.1, HappyHorse 1.0 and SkyReels V4 add synchronized sound to the animated image, so one request returns a clip ready for social or ads.
Image to Video API FAQ
What is an image to video API?
An image to video API animates a still image into a video clip. You send the image URL with a prompt describing motion and camera, and APIMart returns a task ID that you poll for the finished video.
What inputs does an image to video API require?
Usually an image URL plus a text prompt; many models also accept a last frame or extra reference images. Supported image counts, formats, durations and resolutions are listed on each model page.
Which image to video API model should I choose?
Use Wan 3.0, Seedance 2.5 or Kling V3 when you need frame control and consistency, Veo 3.1 or HappyHorse 1.0 for cinematic clips with sound, and Kling 3.0 Turbo for quick drafts.
How does image to video API pricing work?
Image to video is billed per second or per clip depending on the model, resolution and duration, and each model page shows its live pricing table. Every model is 20% off list price, and membership saves up to 28%.
How do I get an image to video API key?
Sign up for APIMart, open the API Keys page in your dashboard and create a key. One key covers every image to video model here plus the rest of the 190+ models.
Is the image to video API free?
No. APIMart has no free credits; billing is pay as you go with no subscription, and failed generations are not charged. Membership discounts range from 20% to 28%.