Happy Horse 1.1 Reference to Video
Generate 3-15 second videos from up to 9 reference images with Alibaba's Happy Horse 1.1. Cite each reference positionally in the prompt — [Image 1], [Image 2] — to lock subjects, products, or styles into the scene, with native audio and nine aspect ratios from 9:21 to 21:9.
23% less than the priciest provider (5 compared). Same model, same output.
1-9 reference images for subjects, products or style. Address them in the prompt as [Image 1], [Image 2].
Duration
Length of the generated video in seconds (3-15). Billed per second.
Resolution
Output resolution. 1080p is the model's default; 480p is the budget rung.
Aspect Ratio
Aspect ratio of the generated video.
23% less than the priciest provider (5 compared). Same model, same output.
Your result renders here.
Use Happy Horse 1.1 Reference to Video via the API
Generate with Happy Horse 1.1 Reference to Video from your own code. Async by design: submit a generation, get a run id back instantly, poll until it completes. Billed in credits at the best available provider price — from $0.07 per second, audio included.
1 — Start a generation
curl -X POST https://app.scenetra.com/api/v1/generate/happy-horse-1.1-ref2v \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"prompt":"The woman in [Image 1] wearing the jacket from [Image 2] walks through a rainy street at night","images":["https://example.com/person.jpg","https://example.com/jacket.jpg"],"duration":"5"}'Response (202)
{
"runId": "run_abc123",
"status": "PENDING",
"model": "happy-horse-1.1-ref2v",
"estimatedCredits": 725,
"statusUrl": "/api/v1/runs/run_abc123"
}2 — Poll until COMPLETED (every 2-5 seconds)
curl https://app.scenetra.com/api/v1/runs/run_abc123 \
-H "Authorization: Bearer YOUR_API_KEY"Response
{
"runId": "run_abc123",
"status": "COMPLETED",
"progress": { "completed": 1, "total": 1 },
"outputs": [
{ "type": "video", "url": "https://images.scenetra.com/..." }
],
"error": null,
"estimatedCredits": 725
}Request fields
| Field | Type | Description | Default |
|---|---|---|---|
| prompt* | string | Describe the video and reference your images positionally: 'the woman in [Image 1] wearing the jacket from [Image 2]'. | — |
| images | string[] (https URLs) | 1-9 reference images for subjects, products or style. Address them in the prompt as [Image 1], [Image 2]. | — |
| duration | string | Length of the generated video in seconds (3-15). Billed per second.3 · 4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 | 5 |
| resolution | string | Output resolution. 1080p is the model's default; 480p is the budget rung.480p · 720p · 1080p | 1080p |
| aspect_ratio | string | Aspect ratio of the generated video.16:9 · 9:16 · 1:1 · 4:3 · 3:4 · 4:5 · 5:4 · 9:21 · 21:9 | 16:9 |
| callbackUrl | string (https URL) | Optional. When the run finishes (completed or failed), we POST the final run status to this URL — same shape as the poll response, plus model, params, and chargedCredits. Best-effort; polling remains the source of truth. | |
Unknown fields are rejected with a helpful 400. Media fields take https URLs. There is no provider parameter — every run is routed to the cheapest capable provider automatically.
Good to know
Billing
Runs bill your account credits at provider cost with 0% markup. The submit response includes the credit estimate; failed generations are refunded automatically. Add ?dryRun=1 to get the estimate without running anything.
Errors
Errors are always { "error": { "code", "message" } } with stable codes: unauthorized, model_not_available, invalid_parameter, insufficient_credits, rate_limited. A failed run is not an HTTP error — poll responses report status: "FAILED" with the reason.
What makes Happy Horse 1.1 Reference to Video different
Native audio, directed by your prompt
Happy Horse 1.1 generates picture and sound together. Dialogue, background music, and sound effects come out of the same prompt as the visuals — write "she whispers 'follow me' over distant thunder" and the model stages both the shot and the soundtrack. There is no separate audio pass and no surcharge: the per-second rate covers everything.
Nine aspect ratios, 9:21 to 21:9
Most video models offer three or four ratios. Happy Horse 1.1 offers nine — 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, and the unusually tall 9:21 and ultrawide 21:9. That means native output for vertical stories, square feeds, classic TV framing, and anamorphic-style cinematics without cropping in post.
Any duration from 3 to 15 seconds
Pick any whole second from 3 to 15. Three-second clips are cheap motion tests — a 3-second 480p draft costs about $0.21 — while 15 seconds is enough for a full beat of action with dialogue. Billing is strictly per second, so an 11-second clip costs exactly 11 seconds, not a rounded-up preset.
Draft at 480p, ship at 1080p
480p costs $0.07 per second, 720p $0.099, and 1080p $0.128. A practical loop on a Scenetra board: iterate your prompt at 480p until the motion and audio land, then flip the resolution select to 1080p for the final render — same node, same wiring, billed at provider cost with 0% markup.
Parameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Describe the scene, action, camera, and style. Audio — dialogue, music, SFX — is generated natively and directed through this prompt. | — |
| Duration | select | Duration of the video in seconds, any whole second from 3 to 15.3 · 4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 | 5 |
| Resolution | select | Output resolution. 1080p is the model's default; 480p is the budget rung.480p · 720p · 1080p | 1080p |
| Aspect Ratio | select | Aspect ratio of the generated video, from tall 9:21 to ultrawide 21:9.16:9 · 9:16 · 1:1 · 4:3 · 3:4 · 4:5 · 5:4 · 9:21 · 21:9 | 16:9 |
Pricing
From $0.07 per second, audio included
Pay per generation — only pay for what you use.
480p
$0.07/s
per second
720p
$0.099/s
per second
1080p
$0.128/s
per second
What a video costs
| Duration | 480p | 720p | 1080p |
|---|---|---|---|
| 5 seconds | $0.35 | $0.49 | $0.64 |
| 10 seconds | $0.70 | $0.99 | $1.28 |
Use cases
Social content
- Vertical 9:16 story clips
- Sound-on hooks with dialogue
- Square feed loops
- Trend-format skits
Film & storytelling
- Ultrawide 21:9 establishing shots
- Dialogue scenes with native audio
- Mood pieces with ambience
- 15-second narrative beats
Marketing & ads
- Product teasers with music
- Multi-ratio campaign variants
- Spoken-tagline bumpers
- Cheap 480p concept tests
Music & audio-led
- Music-driven visuals
- SFX-heavy action clips
- Ambient soundscape loops
- Rhythmic montage drafts
Frequently asked questions
What is reference-to-video in Happy Horse 1.1?+
It's a generation mode that conditions the video on up to 9 reference images instead of a single first frame. You describe the scene in the prompt and cite each reference positionally — [Image 1], [Image 2] — so the model knows which image supplies the character, which the product, and which the style. The output is a 3-15 second clip with native audio.
How do I reference images in the prompt?+
Address them by position in square brackets: "the man in [Image 1] holds the bottle from [Image 2] in a neon-lit bar". Upload order defines the numbering. Being explicit about the role of each image — who is a subject, what is a prop, which sets the style — is what makes multi-reference compositions come out coherent.
What are the requirements for reference images?+
You can use 1 to 9 images in JPEG, PNG, or WEBP, each up to 20MB with the shortest side at least 400px. Clean, well-lit references with the subject clearly visible give the model the most to lock onto.
Can I keep a character consistent across multiple videos?+
Yes — that's the point of this mode. Reuse the same reference images across generations and only change the prompt: the character's face, outfit, or your product stays consistent while the scene, action, and camera change. On a Scenetra board you can duplicate the node and batch out a whole series.
How is this different from image-to-video?+
Image to Video uses one image as the literal first frame and inherits its aspect ratio. Reference to Video uses up to 9 images as identity and style anchors — the video doesn't start from any of them, and you keep full control of aspect ratio (nine options from 9:21 to 21:9). Use image-to-video to animate a finished composition, reference-to-video to cast subjects into new scenes.
Does reference mode also generate audio?+
Yes — native audio comes with every Happy Horse 1.1 mode. Dialogue, music, and sound effects are generated together with the picture and directed through the same prompt, at no extra cost on top of the per-second rate.
Related models
MiniMax H3
MiniMax H3 is the newest generation of the Hailuo video lineage: 5-15 second clips with native stereo audio — dialogue, music, and sound effects rendered in the same pass as the picture. Generate at 480P or 768P natively, or step up to 2K and 4K upscaled output, in six aspect ratios from cinematic 21:9 to vertical 9:16.
View model →Seedance 2.5
ByteDance's Seedance 2.5 generates cinematic videos up to 30 seconds from text prompts with synchronized audio — voice, sound effects, and background music — built in by default. The newest Seedance generation, with mp4 or high-fidelity mov output for multi-round editing pipelines.
View model →PixVerse v6
PixVerse's flagship video model generates 1-15 second clips from a text prompt with optional synchronized audio, multi-shot sequences, four resolutions from 360p to 1080p, and eight aspect ratios from vertical 9:16 to cinematic 21:9. Starting at $0.02 per second, it is one of the most affordable ways to produce finished AI video.
View model →Kling 3.0 Pro
Kuaishou's flagship video generation model delivering stunning visual fidelity. Kling 3.0 Pro supports text-to-video, image-to-video with start/end frames, element references for character consistency, and native audio generation.
View model →Start creating with Happy Horse 1.1 Reference to Video
Use Happy Horse 1.1 Reference to Video alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started