Multi-Character Reference Video with Happy Horse
Alibaba's Happy Horse reference-to-video generates videos from 1-9 reference images. Reference each subject in your prompt as character1, character2, ... character9 (order matches your uploaded images) for consistent multi-character scenes. 720p or 1080p output, 3-15 second durations.
What makes Happy Horse Reference-to-Video different
Up to 9 reference images in one scene
Upload between 1 and 9 reference images and Happy Horse weaves them into a single coherent clip. That's enough for a full cast: a lead character, a supporting character, a product, and a style plate can all anchor the same generation, keeping every subject recognizable across every frame.
Explicit character1-character9 prompt syntax
You don't hope the model guesses which image is which — you tell it. Each uploaded image maps to a slot in upload order, and you direct them in the prompt as character1, character2, and so on: 'character1 hands the box to character2'. That explicit binding is what makes multi-subject scenes directable rather than a lottery.
Full-length, full-resolution reference generations
Reference mode isn't a cut-down variant: you keep the same 720p and 1080p tiers, the same five aspect ratios, and the same 3-15 second per-second duration control as the main Happy Horse model, at the same per-second price.
Character pipelines on one board
On Scenetra you can generate a character with an image model, refine it with an editing node, then wire the results straight into a Happy Horse Reference to Video node — and reuse the same reference set across many shots. That's how you build a recurring character or AI influencer without re-describing them every time.
Playground
A timelapse of a flower blooming in a sunlit meadow, cinematic quality
Drop images or click to upload
Try Happy Horse Reference-to-Video in Scenetra
Open PlaygroundParameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Reference subjects as character1, character2, etc. — order matches the uploaded reference images. Max 2500 characters. | — |
| Reference Images* | image upload | 1-9 reference images for subject consistency. Min 400px shortest side (720p+ recommended), max 10MB each. JPEG/PNG/WEBP. | — |
| Aspect Ratio | select | Aspect ratio of the generated video.16:9 · 9:16 · 1:1 · 4:3 · 3:4 | 16:9 |
| Resolution | select | Output video resolution tier.720p · 1080p | 1080p |
| Duration | select | Output video duration in seconds (3-15).3 · 4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 | 5 |
Pricing
From $0.14 per second
Pay per generation — only pay for what you use.
720p
$0.14/s
per second
1080p
$0.28/s
per second
What a video costs
| Duration | 720p | 1080p |
|---|---|---|
| 5 seconds | $0.70 | $1.40 |
| 10 seconds | $1.40 | $2.80 |
Use cases
Characters & influencers
- Recurring AI characters
- Virtual influencer clips
- Multi-character dialogue scenes
- Character turnaround shots
Brand & product
- Product held consistent across shots
- Mascot animations
- Brand-style locked campaigns
- Packaging in lifestyle scenes
Film & storyboarding
- Cast consistency across previz shots
- Costume and prop continuity
- Scene variations with a fixed cast
- Style-matched sequences
Social series
- Episodic content with one cast
- Vertical 9:16 character skits
- Before/after with the same subject
- Fan content from character sheets
Related models
Seedance 2.5 Reference-to-Video
Generate videos from up to 30 reference images, 10 reference videos, and 10 audio tracks with Seedance 2.5 by ByteDance. Cite references directly in your prompt as @Image1, @Video1, @Audio1. Unique to 2.5: audio-only referencing — a single music or voice track can drive visual pacing, beat matching, and lip-sync.
View model →Seedance 2 Reference-to-Video
Generate videos from up to 9 reference images, 3 reference videos, and 3 audio tracks with Seedance 2.0 by ByteDance. Mention references in your prompt — 'image 1', 'video 1' — for character consistency, motion transfer, and audio-driven lip-sync, at resolutions up to 1080p.
View model →Seedance 2 Mini Reference-to-Video
Seedance 2 Mini Reference-to-Video generates consistent videos from up to 9 reference images, 3 reference videos, and 3 audio tracks. Reference subjects directly in your prompt as 'image 1', 'video 1', etc. for character consistency and motion transfer — at ByteDance's fast, affordable Mini pricing.
View model →Grok Reference-to-Video
Generate videos with consistent subject appearance from up to 7 reference images using xAI's Grok Imagine Video. Reference each image in your prompt as @Image1, @Image2, etc. for style and content guidance.
View model →Frequently asked questions
What is Happy Horse Reference to Video?+
It's the reference-guided mode of Alibaba's Happy Horse video model: instead of starting from text alone or a single first frame, you upload 1-9 reference images and the model keeps those characters, products, and styles consistent throughout the generated clip. On Scenetra it runs as its own node in the visual workflow editor.
How do I reference characters in the prompt?+
Each uploaded image maps to a numbered slot in upload order: the first image is character1, the second is character2, up to character9. Write your prompt using those names — for example, 'character1 walks into frame and greets character2' — and the model binds each subject to the right reference image.
How many reference images can I use, and what are the requirements?+
Between 1 and 9 images per generation, in JPEG, PNG, or WEBP. Each image needs at least 400 pixels on its shortest side (720p or larger is recommended for best identity retention) and must be under 10MB.
How much does Happy Horse reference to video cost?+
The same per-second rate as the main Happy Horse model: $0.14 per second at 720p and $0.28 per second at 1080p, regardless of how many reference images you attach. A 5-second 1080p clip is $1.40; 10 seconds is $2.80.
Can I try Happy Horse reference to video for free?+
You can try it free online: Scenetra runs in the browser, and new accounts get free welcome credits plus a 7-day trial that covers your first generations. After that it's pay-per-generation with no subscription required.
When should I use Reference to Video instead of image to video?+
Use image to video (on the main Happy Horse page) when you want the video to literally start from a specific frame. Use Reference to Video when you care about identity rather than the opening composition — the model composes a new scene freely while keeping your referenced subjects consistent, which is what you want for recurring characters and multi-subject shots.
Start creating with Happy Horse Reference-to-Video
Use Happy Horse Reference-to-Video alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started Free