HappyHorse 是阿里巴巴 ATH-AI 推出的统一多模态视频生成模型,单次前向传播即可同步生成 1080p 视频与原生音频,支持 7 语言对口型,在 Artificial Analysis 文生视频与图生视频双榜中均排名第一。
Video link valid for 72 hours
透明定价,无隐藏费用。按量付费,用多少付多少。
* 实际费用以最终输出为准。
阿里巴巴 HappyHorse 在同一个 Transformer 中统一画面与音频 token,输出 1080p 视频,支持 7 种语言的亚像素级对口型。APIMart 提供统一的 HappyHorse API 接入,仅需一个 API Key。
50K+
活跃用户
99.9%
在线率
2x
更快
70%
成本节省
首个原生协同生成画面与音频的视频模型
受益于原生音画协同生成的生产级场景
几步即可开始创作原生音频 AI 视频
免费注册 APIMart 账号,开始使用 HappyHorse。任何时候都可以为你的团队创建组织。
为账户充值后即可开始使用,余额可在平台所有模型间通用,包括 HappyHorse。
在控制台创建 API Key,调用 HappyHorse 时用于身份验证,立即获得原生音频 AI 视频能力。
全球创作者与开发者的真实反馈
“原生音频效果太惊艳了 —— 一次出片就能拿到对得上口型的对白,完全不用再走 TTS 流水线。”
Alex Morgan
创意总监
“HappyHorse 把我们的本地化耗时砍掉了 70%,一个 prompt 输出 7 种语言版本,口型全部对齐。”
Sarah Kim
市场经理
“原生 1080p 出片,没有放大伪影。多镜头序列的时序一致性表现非常稳。”
James Wilson
全栈开发工程师
“把之前的视频供应商换成了 HappyHorse,对白密集型场景的画质与同步度立刻提升。”
Lisa Chen
StartupAI CTO
“作为独立电影人,HappyHorse 让我能在分镜阶段就生成带临时对白的预演镜头,效率提升明显。”
Marco Rivera
独立电影人
“通过 APIMart 的统一 API 调用 HappyHorse,一个 Key 走天下,集成不到一小时就跑通了。”
Emily Zhang
DevOps 工程师
在 APIMart 使用 HappyHorse 你需要了解的一切
HappyHorse 是阿里巴巴 ATH-AI 创新业务部(源自淘宝与天猫未来生活实验室)推出的视频生成模型。它是一个单流多模态 Transformer,可在单次前向传播中同步生成 1080p 视频与原生音频。
HappyHorse 接受以下核心参数:model(happyhorse-1.0 或 happyhorse-1.1)、prompt(文本描述,≤2500 字符)、resolution(720P 或 1080P,默认 1080P)、duration(时长 3-15 秒,默认 5 秒)、aspect_ratio(16:9 / 9:16 / 1:1 / 4:3 / 3:4,仅文生视频生效),可选 seed 和 watermark。传入 image_urls 或 first_frame_image 即切换到图生视频模式(此时 aspect_ratio 由首帧图决定)。
大多数视频模型先生成画面、再单独添加音频。HappyHorse 在同一个 Transformer 中协同生成画面与音频,口型对齐到亚像素级,环境音也天然同步,对白场景表现尤为真实。
HappyHorse 支持 7 种语言:英语、普通话、粤语、日语、韩语、德语、法语。模型会根据所选语种在亚像素精度上对齐口型。
在单张 NVIDIA H100 GPU 上,HappyHorse 生成一段原生 1080p 视频约需 38 秒。它采用 8 步 DMD-2 采样、跳过 classifier-free guidance,比常规 30+ 步扩散视频模型显著更快。
在 Artificial Analysis 盲评中,HappyHorse 文生视频(1333 Elo)与图生视频(1392 Elo)均排名第一。与 Sora 2、Veo 3.1 相比,最大差异是原生音频协同生成;与可灵(Kling)相比,对白与口型表现更优,但可灵在部分动作密集场景仍具优势。
请参见本页价格区域查看最新的按秒计费价格。在 APIMart 充值后即可开始使用 HappyHorse 生成视频。
向 /v1/videos/generations 发送 POST 请求,携带 APIMart API Key、模型名 happyhorse-1.0 或 happyhorse-1.1 与 prompt 即可。完整集成示例见 API 文档。
探索同类型的其他模型。

SkyReels V4 Fast
SkyReels-V4-Fast is the first unified framework for joint audiovisual generation, restoration, and editing at cinema-grade quality, efficiently achieving 1080p, 15-second multi-shot video generation with synchronized audio and visuals through a dual-stream MMDiT architecture

Wan 2.7
Wan-2.7 is Alibaba’s next-generation multimodal AI video model that generates high-quality videos from text prompts, images, or reference footage. It supports text-to-video, image-to-video, and instruction-based video editing, producing short clips (up to ~15 seconds) with 720p–1080p resolution, realistic motion, and strong character consistency. 

ViduQ 3
Vidu Q3 is an advanced AI video generation model developed by Shengshu Technology that creates cinematic videos from text prompts or images. It supports both text-to-video and image-to-video workflows, generating clips up to around 16 seconds with synchronized native audio, including dialogue and sound effects.

Seedance 2.0
seedance-2-0 (Seedance 2.0) is the second-generation multimodal audio-video generation large model launched by ByteDance. It supports the fusion of various inputs such as text, images, audio, and video, enabling the efficient creation of high-quality video content. Its core technologies include multimodal joint generation, physical logic optimization, and audiovisual integration. It is widely used in industries such as film, advertising, and e-commerce, providing creators with powerful video editing and creation tools, supporting natural scene continuation and dynamic adjustments. The model outperforms competitors in multiple technical aspects and represents a significant breakthrough in AI video generation.