Cinematic AI Video with Alibaba Happy Horse
Alibaba's Happy Horse generates cinematic videos from text prompts or animates a first-frame image into a full video. 1080p output, durations 3-15 seconds, support for both text-to-video and image-to-video workflows.
What makes Happy Horse different
Any duration from 3 to 15 seconds
Happy Horse lets you pick any whole-second duration between 3 and 15 seconds — not just the fixed 5-or-10 choice most video models offer. Because billing is per second, a 3-second sting costs a fraction of a full clip, which makes it practical to generate short cutaways, transitions, and loops without paying for footage you'll trim away.
Text to video and image to video in one node
Write a prompt of up to 2500 characters and Happy Horse invents the scene from scratch. Drop an optional first-frame image into the same node and it switches to image-to-video, animating your exact starting composition instead. One node on your Scenetra board covers both workflows — no swapping models when you decide to start from a still.
Cinematic 1080p output
Happy Horse renders at 720p or full 1080p, with five aspect ratios covering widescreen 16:9, vertical 9:16 for shorts and stories, square 1:1, and the 4:3 / 3:4 pair. Alibaba positioned the model squarely at cinematic look and camera movement, and the 1080p tier is where that shows.
A second mode for character consistency
The model on this page starts from text or a single first frame. When you need the same character, product, or style held consistent across shots, switch to Happy Horse Reference to Video — it accepts 1-9 reference images and lets you cite each one in the prompt as character1 through character9. Both modes sit side by side on the same Scenetra board.
Playground
A timelapse of a flower blooming in a sunlit meadow, cinematic quality
Drop images or click to upload
Try Happy Horse in Scenetra
Open PlaygroundParameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Text prompt describing the desired video (max 2500 characters). | — |
| First frame | image upload | Optional first-frame image to enable image-to-video mode. | — |
| Aspect Ratio | select | Aspect ratio of the generated video.16:9 · 9:16 · 1:1 · 4:3 · 3:4 | 16:9 |
| Resolution | select | Output video resolution.720p · 1080p | 1080p |
| Duration | select | Output video duration in seconds (3-15).3 · 4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 | 5 |
Pricing
From $0.14 per second
Pay per generation — only pay for what you use.
720p
$0.14/s
per second
1080p
$0.28/s
per second
What a video costs
| Duration | 720p | 1080p |
|---|---|---|
| 5 seconds | $0.70 | $1.40 |
| 10 seconds | $1.40 | $2.80 |
Use cases
Marketing & ads
- Product spots from a prompt
- Concept testing before a shoot
- Brand stings and bumpers
- Localized ad variants
Social content
- Vertical 9:16 shorts
- Cinematic b-roll
- Hook clips in odd durations
- Story sequences
Film & previz
- Establishing shots
- Animatics from key frames
- Camera-move studies
- Mood and tone tests
E-commerce
- Animate product photos
- Lifestyle scene clips
- Seasonal campaign refreshes
- Listing videos from stills
Related models
Wan 2.6
A versatile text-to-video and image-to-video model with built-in audio support and prompt expansion. Wan 2.6 delivers up to 1080p resolution with 15-second videos, making it a solid all-rounder for creative video generation.
View model →Seedance 2
ByteDance's Seedance 2.0 generates cinematic videos from text prompts with synchronized audio — voice, sound effects, and background music. Pick any duration from 4 to 15 seconds and resolutions up to 1080p, with per-second pricing that keeps 480p drafting cheap.
View model →Kling v2.6
Kling v2.6 generates videos from text prompts or first-frame images with native synchronized audio support, including dialogue in Chinese and English. Built for short-form social and cinematic content.
View model →Hailuo 2.3
MiniMax's Hailuo 2.3 video generation model delivering 768p videos with both text-to-video and image-to-video support. A cost-effective option for social media content, quick iterations, and template-based video creation.
View model →Frequently asked questions
What is the Happy Horse video model?+
Happy Horse is Alibaba's AI video generation model. It creates cinematic clips from text prompts, or animates a first-frame image into a video, at 720p or 1080p and any duration from 3 to 15 seconds. On Scenetra it runs as a node in the visual workflow editor, alongside a separate Reference to Video mode for character consistency.
How can I use Happy Horse?+
The fastest way is in the browser: sign up for Scenetra, drop a Happy Horse node onto your board, write a prompt (optionally attach a first-frame image), and generate. No API key, no waitlist, no local GPU — Scenetra is web-based and works anywhere, and Happy Horse sits next to every other major video model on the same canvas.
Is Happy Horse free or paid?+
It's a paid model billed per second of video, but you can try it without paying: new Scenetra accounts start with free welcome credits and a 7-day trial, which covers your first generations. After that you pay per generation from your credit balance — there is no subscription gating the model and nothing to cancel.
How much does Happy Horse cost?+
On Scenetra, Happy Horse costs $0.14 per second at 720p and $0.28 per second at 1080p. A 5-second 1080p clip is $1.40; a 10-second one is $2.80. Because duration is adjustable per second from 3 to 15, you only pay for exactly the length you generate.
Is Happy Horse better than Seedance?+
They trade blows. Early comparisons generally rate Seedance 2.0 as the more established, more consistent model, while Happy Horse is newer and competes hard on cinematic look and price. On Scenetra you don't have to pick in advance — both are available on the same board, so you can run the same prompt through each and judge on your own footage.
Can Happy Horse do image to video?+
Yes. Attach a first-frame image to the node and Happy Horse animates from that exact composition, guided by your prompt. If you instead need one or more subjects held consistent across a scene, use the dedicated Reference to Video mode, which accepts 1-9 reference images.
What resolutions, durations, and aspect ratios does Happy Horse support?+
720p and 1080p output; any whole-second duration from 3 to 15 seconds; and aspect ratios 16:9, 9:16, 1:1, 4:3, and 3:4. Prompts can be up to 2500 characters.
Is Happy Horse open source?+
No — Alibaba has not released open weights for Happy Horse; it's available through hosted APIs. On Scenetra you can use it directly in the browser with no API setup, and compare it side by side with open-weight models like LTX-2 and Wan on the same board.
Start creating with Happy Horse
Use Happy Horse alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started Free