HappyHorse is Alibaba ATH-AI's unified multimodal video model, available on APIMart as the HappyHorse API. It generates 1080p video with synchronized native audio in a single forward pass, supports 7-language lip-sync, and ranks #1 on Artificial Analysis text-to-video and image-to-video leaderboards.
Video link valid for 72 hours
Transparent pricing with no hidden fees. Pay only for what you use.
* Actual costs are subject to final output.
Alibaba's HappyHorse unifies pixels and audio in a single Transformer, delivering 1080p videos with sub-pixel lip-sync across 7 languages. APIMart provides unified access to the HappyHorse API with a single key.
50K+
Active Users
99.9%
Uptime
2x
Faster
70%
Cost Savings
HappyHorse is the first generative video model to natively co-produce visuals and audio
Production scenarios that benefit from native audio-video co-generation
Start creating native-audio AI videos with HappyHorse in just a few simple steps
Create your free APIMart account to get started with HappyHorse. You can set up an organization for your team at any time.
Add funds to your account balance to start using the service. Your balance can be used across all models on APIMart, including the HappyHorse API.
Create an API key in the dashboard — you'll need it to authenticate every call to HappyHorse. Get instant access to native-audio AI video.
Real feedback from creators and developers worldwide
“The HappyHorse API audio output is unreal — dialogue lip-sync just works on the first generation, no separate TTS pipeline.”
Alex Morgan
Creative Director
“HappyHorse cut our localization time by 70%. One prompt, seven languages, all with matching mouth shapes.”
Sarah Kim
Marketing Manager
“1080p straight out of HappyHorse with no upscaling artifacts. The temporal consistency across multi-shot sequences is impressive.”
James Wilson
Full-Stack Developer
“We swapped our previous video provider for HappyHorse and saw immediate quality wins on dialogue-heavy scenes.”
Lisa Chen
CTO, StartupAI
“As an indie filmmaker, HappyHorse has been a game-changer for storyboard pre-visualization with temp dialogue baked in.”
Marco Rivera
Independent Filmmaker
“Routing the HappyHorse API through APIMart's unified gateway means I keep one key for everything. Integration took less than an hour.”
Emily Zhang
DevOps Engineer
Everything you need to know about the HappyHorse API on APIMart
The HappyHorse API is a generative video service from Alibaba's ATH-AI division (originating from Taobao & Tmall's Future Lifestyle Lab). It exposes a single-stream multimodal Transformer that generates 1080p video and synchronized native audio in one forward pass.
The HappyHorse API accepts these core parameters: model (happyhorse-1.0), prompt (text description, ≤2500 chars), resolution (720P or 1080P, default 1080P), duration (clip length 3-15s, default 5s), aspect_ratio (16:9 / 9:16 / 1:1 / 4:3 / 3:4, text-to-video only), plus optional seed and watermark. Pass image_urls or first_frame_image to switch into image-to-video mode (aspect_ratio is then derived from the first frame).
Most video models generate visuals first, then add audio in a separate step. HappyHorse generates visuals and audio jointly in the same Transformer — yielding sub-pixel lip-sync, naturally aligned ambient sound, and far more believable dialogue scenes.
HappyHorse supports seven languages: English, Mandarin Chinese, Cantonese, Japanese, Korean, German, and French. Mouth shapes align at sub-pixel precision based on the chosen audio language.
On a single NVIDIA H100 GPU, HappyHorse generates a native 1080p clip in roughly 38 seconds. The 8-step DMD-2 sampler skips classifier-free guidance, making it dramatically faster than diffusion video models that require 30+ steps.
On Artificial Analysis blind evaluations, HappyHorse ranks #1 on both text-to-video (1333 Elo) and image-to-video (1392 Elo). Its key differentiator vs. Sora 2 and Veo 3.1 is native audio co-generation; vs. Kling, it offers stronger dialogue and lip-sync.
Check the pricing section on this page for current per-second rates. Top up credits on APIMart to start generating videos with HappyHorse.
Send a POST request to /v1/videos/generations with your APIMart API key, model name happyhorse-1.0, and your prompt. See the HappyHorse documentation for full integration examples.
Explore more models in the same category.

SkyReels V4 Fast
SkyReels-V4-Fast is the first unified framework for joint audiovisual generation, restoration, and editing at cinema-grade quality, efficiently achieving 1080p, 15-second multi-shot video generation with synchronized audio and visuals through a dual-stream MMDiT architecture

Wan 2.7
Wan-2.7 is Alibaba’s next-generation multimodal AI video model that generates high-quality videos from text prompts, images, or reference footage. It supports text-to-video, image-to-video, and instruction-based video editing, producing short clips (up to ~15 seconds) with 720p–1080p resolution, realistic motion, and strong character consistency. 

ViduQ 3
Vidu Q3 is an advanced AI video generation model developed by Shengshu Technology that creates cinematic videos from text prompts or images. It supports both text-to-video and image-to-video workflows, generating clips up to around 16 seconds with synchronized native audio, including dialogue and sound effects.

Doubao Seedance 2.0
doubao-seedance-2-0 (Seedance 2.0) is the second-generation multimodal audio-video generation large model launched by ByteDance. It supports the fusion of various inputs such as text, images, audio, and video, enabling the efficient creation of high-quality video content. Its core technologies include multimodal joint generation, physical logic optimization, and audiovisual integration. It is widely used in industries such as film, advertising, and e-commerce, providing creators with powerful video editing and creation tools, supporting natural scene continuation and dynamic adjustments. The model outperforms competitors in multiple technical aspects and represents a significant breakthrough in AI video generation.