Text & Image to VideoReference to Video

Multi-Character Reference Video with Happy Horse

Alibaba's Happy Horse reference-to-video generates videos from 1-9 reference images. Reference each subject in your prompt as character1, character2, ... character9 (order matches your uploaded images) for consistent multi-character scenes. 720p or 1080p output, 3-15 second durations.

What makes Happy Horse Reference-to-Video different

Up to 9 reference images in one scene

Upload between 1 and 9 reference images and Happy Horse weaves them into a single coherent clip. That's enough for a full cast: a lead character, a supporting character, a product, and a style plate can all anchor the same generation, keeping every subject recognizable across every frame.

Explicit character1-character9 prompt syntax

You don't hope the model guesses which image is which — you tell it. Each uploaded image maps to a slot in upload order, and you direct them in the prompt as character1, character2, and so on: 'character1 hands the box to character2'. That explicit binding is what makes multi-subject scenes directable rather than a lottery.

Full-length, full-resolution reference generations

Reference mode isn't a cut-down variant: you keep the same 720p and 1080p tiers, the same five aspect ratios, and the same 3-15 second per-second duration control as the main Happy Horse model, at the same per-second price.

Character pipelines on one board

On Scenetra you can generate a character with an image model, refine it with an editing node, then wire the results straight into a Happy Horse Reference to Video node — and reuse the same reference set across many shots. That's how you build a recurring character or AI influencer without re-describing them every time.

Playground

Happy Horse Reference-to-Video
playground

A timelapse of a flower blooming in a sunlit meadow, cinematic quality

Drop images or click to upload

16:9
1080p
5
Generate
Resultpreview
Sample output from Happy Horse Reference-to-Video

Try Happy Horse Reference-to-Video in Scenetra

Open Playground

Parameters

ParameterTypeDescriptionDefault
Prompt*textReference subjects as character1, character2, etc. — order matches the uploaded reference images. Max 2500 characters.
Reference Images*image upload1-9 reference images for subject consistency. Min 400px shortest side (720p+ recommended), max 10MB each. JPEG/PNG/WEBP.
Aspect RatioselectAspect ratio of the generated video.16:9 · 9:16 · 1:1 · 4:3 · 3:416:9
ResolutionselectOutput video resolution tier.720p · 1080p1080p
DurationselectOutput video duration in seconds (3-15).3 · 4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 155

Pricing

From $0.14 per second

Pay per generation — only pay for what you use.

720p

$0.14/s

per second

1080p

$0.28/s

per second

What a video costs

Duration720p1080p
5 seconds$0.70$1.40
10 seconds$1.40$2.80

Use cases

Characters & influencers

  • Recurring AI characters
  • Virtual influencer clips
  • Multi-character dialogue scenes
  • Character turnaround shots

Brand & product

  • Product held consistent across shots
  • Mascot animations
  • Brand-style locked campaigns
  • Packaging in lifestyle scenes

Film & storyboarding

  • Cast consistency across previz shots
  • Costume and prop continuity
  • Scene variations with a fixed cast
  • Style-matched sequences

Social series

  • Episodic content with one cast
  • Vertical 9:16 character skits
  • Before/after with the same subject
  • Fan content from character sheets

Related models

Frequently asked questions

What is Happy Horse Reference to Video?+

It's the reference-guided mode of Alibaba's Happy Horse video model: instead of starting from text alone or a single first frame, you upload 1-9 reference images and the model keeps those characters, products, and styles consistent throughout the generated clip. On Scenetra it runs as its own node in the visual workflow editor.

How do I reference characters in the prompt?+

Each uploaded image maps to a numbered slot in upload order: the first image is character1, the second is character2, up to character9. Write your prompt using those names — for example, 'character1 walks into frame and greets character2' — and the model binds each subject to the right reference image.

How many reference images can I use, and what are the requirements?+

Between 1 and 9 images per generation, in JPEG, PNG, or WEBP. Each image needs at least 400 pixels on its shortest side (720p or larger is recommended for best identity retention) and must be under 10MB.

How much does Happy Horse reference to video cost?+

The same per-second rate as the main Happy Horse model: $0.14 per second at 720p and $0.28 per second at 1080p, regardless of how many reference images you attach. A 5-second 1080p clip is $1.40; 10 seconds is $2.80.

Can I try Happy Horse reference to video for free?+

You can try it free online: Scenetra runs in the browser, and new accounts get free welcome credits plus a 7-day trial that covers your first generations. After that it's pay-per-generation with no subscription required.

When should I use Reference to Video instead of image to video?+

Use image to video (on the main Happy Horse page) when you want the video to literally start from a specific frame. Use Reference to Video when you care about identity rather than the opening composition — the model composes a new scene freely while keeping your referenced subjects consistent, which is what you want for recurring characters and multi-subject shots.

Start creating with Happy Horse Reference-to-Video

Use Happy Horse Reference-to-Video alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.

Get Started Free