kling-motion-control 是 APIMart 上的 Kling 运动控制模型。它需要参考图片与参考视频,可选提示词,并支持角色朝向、生成模式、原声保留和水印参数。
Video link valid for 72 hours
透明定价,无隐藏费用。按量付费,用多少付多少。
* 实际费用以最终输出为准。
在 APIMart 访问 kling-motion-control,上传参考图片和参考视频,使用 Kling 生成可控运动视频。
50K+
活跃用户
99.9%
在线率
2x
更快
70%
成本节省
让 Kling 运动控制脱颖而出的核心特性
由参考图片与参考视频驱动的 Kling 运动控制在真实业务中的应用
三步生成由参考素材驱动的 Kling 运动视频
在 APIMart 注册并创建 API 密钥,即可访问 kling-motion-control。
上传参考图片和参考视频,选择角色朝向、模式、原声和水印选项,然后提交任务。
使用 motion-control 端点传入 image_url、video_url、character_orientation、mode,并按需添加 prompt、element_list 等可选字段。
团队对 Kling 运动控制的真实反馈
“kling-motion-control 正是我们需要的快速迭代工具。参考图锁定主体,参考视频带来动作节奏,输出稳定很多。”
Sarah Johnson
创意总监
“我们把 kling-motion-control 接入 Kling 流水线后立刻缩短了集成时间。极简的 API 让规模化变得轻松。”
James Liu
高级工程师
关于 Kling 运动控制的常见问题
kling-motion-control 专注于参考素材驱动的运动控制。Playground 暴露参考图片、参考视频、角色朝向、模式、原声和水印等字段,与当前 API 参数保持一致。
需要。当前接口要求同时提供 image_url 和 video_url,prompt 是可选字段,可用于补充生成意图。
参考视频至少 3 秒。character_orientation 为 image 时最长 10 秒,为 video 时最长 30 秒。
按秒计费,与 Kling 基础档位相同。当前费率以本页定价区域为准。
可以。向 /kling/v1/videos/motion-control 提交 model、image_url、video_url、character_orientation 和 mode,接口会返回 task ID 供你轮询最终视频 URL。
你可以通过页面右下角的在线客服、发送邮件至 [email protected],或在 Discord 社区中联系我们。我们的团队会尽快回复你的问题。
探索同类型的其他模型。

HappyHorse
HappyHorse-1.0 is an AI model that generates videos based on text input.

SkyReels V4 Fast
SkyReels-V4-Fast is the first unified framework for joint audiovisual generation, restoration, and editing at cinema-grade quality, efficiently achieving 1080p, 15-second multi-shot video generation with synchronized audio and visuals through a dual-stream MMDiT architecture

Wan 2.7
Wan-2.7 is Alibaba’s next-generation multimodal AI video model that generates high-quality videos from text prompts, images, or reference footage. It supports text-to-video, image-to-video, and instruction-based video editing, producing short clips (up to ~15 seconds) with 720p–1080p resolution, realistic motion, and strong character consistency. 

ViduQ 3
Vidu Q3 is an advanced AI video generation model developed by Shengshu Technology that creates cinematic videos from text prompts or images. It supports both text-to-video and image-to-video workflows, generating clips up to around 16 seconds with synchronized native audio, including dialogue and sound effects.