Reference-to-Video with Native Audio using Happy Horse 1.1
Generate 3-15 second videos from up to 9 reference images with Alibaba's Happy Horse 1.1. Cite each reference positionally in the prompt — [Image 1], [Image 2] — to lock subjects, products, or styles into the scene, with native audio and nine aspect ratios from 9:21 to 21:9.
What makes Happy Horse 1.1 different
Up to 9 references, addressed by name
Upload as many as nine reference images — a face, a product, an outfit, a style frame — and cite each one positionally in the prompt: "the woman in [Image 1] wearing the jacket from [Image 2] walks through the market in [Image 3]". The model composes them into one coherent scene instead of guessing which image means what.
Consistent subjects across generations
Because references pin identity, the same character or product stays recognizable from clip to clip. Keep the reference set fixed on your Scenetra board and vary only the prompt to build a series — different scenes, same face, same product — which is exactly what brand and character work needs.
Native audio and full framing control
Like every Happy Horse 1.1 mode, audio is generated natively with the picture and directed through the prompt — no surcharge and no separate pass. Unlike image-to-video, reference mode keeps the full nine-way aspect-ratio choice, from tall 9:21 through ultrawide 21:9, at 480p, 720p, or 1080p.
Reference mode pricing
480p stays at $0.07 per second, while the extra conditioning work prices 720p at $0.113 and 1080p at $0.145 per second — a modest step up from the text and image modes. A 5-second 1080p reference clip is about $0.73, billed at provider cost with 0% markup.
Playground
A timelapse of a flower blooming in a sunlit meadow, cinematic quality
Drop images or click to upload
Try Happy Horse 1.1 in Scenetra
Open PlaygroundParameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Describe the video and reference your images positionally: "the woman in [Image 1] wearing the jacket from [Image 2]". Audio is generated natively. | — |
| Reference Images | image upload | 1-9 reference images for subjects, products, or style. Address them in the prompt positionally: [Image 1], [Image 2], and so on. JPEG/PNG/WEBP, shortest side at least 400px, up to 20MB each. | — |
| Duration | select | Duration of the video in seconds, any whole second from 3 to 15.3 · 4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 | 5 |
| Resolution | select | Output resolution. 1080p is the model's default; 480p is the budget rung.480p · 720p · 1080p | 1080p |
| Aspect Ratio | select | Aspect ratio of the generated video, from tall 9:21 to ultrawide 21:9.16:9 · 9:16 · 1:1 · 4:3 · 3:4 · 4:5 · 5:4 · 9:21 · 21:9 | 16:9 |
Pricing
From $0.07 per second, audio included
Pay per generation — only pay for what you use.
480p
$0.07/s
per second
720p
$0.113/s
per second
1080p
$0.145/s
per second
What a video costs
| Duration | 480p | 720p | 1080p |
|---|---|---|---|
| 5 seconds | $0.35 | $0.56 | $0.73 |
| 10 seconds | $0.70 | $1.13 | $1.45 |
Use cases
Brand & product
- Same product across scene variants
- Model wearing your exact garment
- Packaging in lifestyle settings
- Campaign series with one hero item
Characters & storytelling
- Recurring characters across episodes
- Cast two subjects in one scene
- Style-frame-matched sequences
- Consistent mascot content
Fashion & creators
- Outfit swaps on the same person
- Virtual try-on style clips
- Creator likeness in new settings
- Lookbook videos with audio
Agencies & pitches
- Client assets composed into previz
- Talent references in mood films
- Location plus subject mashups
- Multi-ratio deliverable sets
Related models
MiniMax H3 Reference-to-Video
Direct MiniMax H3 (Hailuo 03) with the material you already have: up to 9 reference images, 3 reference video clips, and 3 audio tracks in a single generation. Keep subjects and products consistent across shots, borrow motion and pacing from reference footage, sync action to a music track, and get native audio in the output — 5-15 seconds, from 480P up to 4K.
View model →Seedance 2.5 Reference-to-Video
Generate videos from up to 30 reference images, 10 reference videos, and 10 audio tracks with Seedance 2.5 by ByteDance. Cite references directly in your prompt as @Image1, @Video1, @Audio1. Unique to 2.5: audio-only referencing — a single music or voice track can drive visual pacing, beat matching, and lip-sync.
View model →PixVerse v6
Drive PixVerse v6 with up to 7 reference images and bind them directly in your prompt as @image1, @image2. Keep characters, objects and scenes consistent across 1-15 second videos at 360p to 1080p, with eight aspect ratios and optional synchronized audio. From $0.0225 per second.
View model →Wan 3.0
Generate video from up to 10 reference images with Alibaba's Wan 3.0. The model carries the people, objects, and styles from your references into a continuous take of 2 to 30 seconds, with native synchronized audio generated in the same pass. Output at 480p, 720p, or 1080p — reference images are free, and audio adds nothing to the price.
View model →Frequently asked questions
What is reference-to-video in Happy Horse 1.1?+
It's a generation mode that conditions the video on up to 9 reference images instead of a single first frame. You describe the scene in the prompt and cite each reference positionally — [Image 1], [Image 2] — so the model knows which image supplies the character, which the product, and which the style. The output is a 3-15 second clip with native audio.
How do I reference images in the prompt?+
Address them by position in square brackets: "the man in [Image 1] holds the bottle from [Image 2] in a neon-lit bar". Upload order defines the numbering. Being explicit about the role of each image — who is a subject, what is a prop, which sets the style — is what makes multi-reference compositions come out coherent.
How much does Happy Horse 1.1 reference-to-video cost?+
$0.07 per second at 480p, $0.113 at 720p, and $0.145 at 1080p, audio included — slightly above the text and image modes at the higher resolutions. A 5-second 1080p clip is about $0.73 and a 10-second one about $1.45, billed at provider cost with 0% markup.
What are the requirements for reference images?+
You can use 1 to 9 images in JPEG, PNG, or WEBP, each up to 20MB with the shortest side at least 400px. Clean, well-lit references with the subject clearly visible give the model the most to lock onto.
Can I keep a character consistent across multiple videos?+
Yes — that's the point of this mode. Reuse the same reference images across generations and only change the prompt: the character's face, outfit, or your product stays consistent while the scene, action, and camera change. On a Scenetra board you can duplicate the node and batch out a whole series.
How is this different from image-to-video?+
Image to Video uses one image as the literal first frame and inherits its aspect ratio. Reference to Video uses up to 9 images as identity and style anchors — the video doesn't start from any of them, and you keep full control of aspect ratio (nine options from 9:21 to 21:9). Use image-to-video to animate a finished composition, reference-to-video to cast subjects into new scenes.
Does reference mode also generate audio?+
Yes — native audio comes with every Happy Horse 1.1 mode. Dialogue, music, and sound effects are generated together with the picture and directed through the same prompt, at no extra cost on top of the per-second rate.
Start creating with Happy Horse 1.1
Use Happy Horse 1.1 alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started Free