Text to VideoImage to VideoReference to Video

Reference-to-Video with Alibaba Wan 3.0

Generate video from up to 10 reference images with Alibaba's Wan 3.0. The model carries the people, objects, and styles from your references into a continuous take of 2 to 30 seconds, with native synchronized audio generated in the same pass. Output at 480p, 720p, or 1080p — reference images are free, and audio adds nothing to the price.

What makes Wan 3.0 different

Up to 10 references, one consistent cast

Feed Wan 3.0 between 1 and 10 reference images and it carries those people, objects, and styles into the generated video. That's the difference between an image-to-video starting frame and true reference conditioning: the references don't have to appear as frame one — they define who and what shows up anywhere in the shot, so a character can walk into a scene that none of your images contain.

30-second takes with native audio

Like the other Wan 3.0 modes, reference-to-video generates any duration from 2 to 30 seconds as one continuous take, with dialogue, sound effects, and music synchronized in the same pass. In this mode audio truly costs nothing extra — the per-second rate is identical whether the toggle is on or off — and the reference images are free at every resolution.

Flat, predictable pricing

Reference-to-video is billed at a flat $0.04 per second at 480p, $0.08 at 720p, and $0.16 at 1080p — higher than the cheapest text- and image-to-video rates, which is the premium for reference conditioning. A 10-second 720p clip with a consistent character is $0.80; a full 30-second 1080p take is $4.80. Draft at 480p to lock the likeness before rendering finals.

Build a character pipeline on one board

On a Scenetra board, generate a character sheet with an image model, wire the outputs into the Wan 3.0 reference node, and produce shot after shot with the same face, wardrobe, and style — a repeatable character pipeline on one canvas. Everything is billed at provider cost with 0% markup, and new accounts get free welcome credits plus a 7-day trial.

Playground

Wan 3.0
playground

A timelapse of a flower blooming in a sunlit meadow, cinematic quality

720p
Auto
5
Generate
Resultpreview
Sample output from Wan 3.0

Try Wan 3.0 in Scenetra

Open Playground

Parameters

ParameterTypeDescriptionDefault
Prompt*textDescribe the video: subject, motion, camera work, and audio direction. Native synchronized audio is generated in the same pass.
Reference Images*image upload1-10 reference images contributing people, objects, or styles to the generated video.
ResolutionselectOutput resolution. 480p is the budget rung; 1080p is the premium rung.480p · 720p · 1080p720p
Aspect RatioselectAspect ratio of the generated video. Auto lets the model choose the best framing for your prompt.Auto · 16:9 · 4:3 · 1:1 · 3:4 · 9:16Auto
DurationselectLength of the generated video in seconds. Wan 3.0 generates any length from 2 to 30 seconds in a single take.2 · 3 · 4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 · 16 · 17 · 18 · 19 · 20 · 21 · 22 · 23 · 24 · 25 · 26 · 27 · 28 · 29 · 305
AudiobooleanGenerate synchronized audio (dialogue, SFX, music) in the same pass. Same price either way in this mode.true

Pricing

From $0.04 per second

Pay per generation — only pay for what you use.

480p

$0.04/s

per second

720p

$0.08/s

per second

1080p

$0.16/s

per second

Audio

Included

per second

Reference images

Free

per second

What a video costs

Duration480p720p1080p
5 seconds$0.20$0.40$0.80
10 seconds$0.40$0.80$1.60

Use cases

Character content

  • Recurring characters across episodes
  • Consistent faces shot after shot
  • Virtual influencer clips
  • Character dialogue scenes with audio

Brand & product

  • Same product in new scenes
  • Brand mascots in motion
  • Style-matched campaign series
  • Wardrobe and prop continuity

Film & previz

  • Cast continuity in previz
  • Scene variations with one look
  • Style-frame-driven sequences
  • 30-second scene blocking

Social series

  • Serialized character shorts
  • Consistent visual identity per feed
  • Multi-part stories, one cast
  • Sound-on episodes up to 30 seconds

Related models

Frequently asked questions

What is Wan 3.0 reference-to-video?+

It's the mode of Alibaba's Wan 3.0 that generates video from up to 10 reference images. Instead of animating one starting frame, the model uses your references to keep specific people, objects, and styles consistent throughout the generated clip — 2 to 30 seconds in a single take, with native synchronized audio. On Scenetra it runs as a node in the visual workflow editor.

How much does Wan 3.0 reference-to-video cost?+

$0.04 per second at 480p, $0.08 at 720p, and $0.16 at 1080p — flat rates, with the reference images free and audio included at no extra cost. A 5-second 720p clip is $0.40 and a 10-second one $0.80. Everything is billed at provider cost with 0% markup.

Why does reference-to-video cost more than Wan 3.0 text-to-video?+

Reference conditioning is billed on a different, flat price ladder. Text- and image-to-video with audio route to a cheaper rate (from $0.025 per second at 480p), while reference-to-video is $0.04 to $0.16 per second depending on resolution regardless of the audio setting. The premium buys consistency: the same faces, objects, and styles across every generation.

How many reference images can I use?+

Between 1 and 10 per generation, and they're free — you pay only the per-second video rate. You can mix reference types in one job: a character, a product, and a style frame together, with the prompt describing how they interact.

What's the difference between reference-to-video and image-to-video?+

Image-to-video animates one upload as the literal first frame of the clip. Reference-to-video accepts up to 10 images that condition the content — the people, objects, and styles they contain appear consistently anywhere in the shot, in scenes your images never showed. Use image-to-video to bring a finished composition to life, and reference-to-video to keep a cast or look consistent across many shots.

Does reference-to-video generate audio too?+

Yes — like the other Wan 3.0 modes it generates dialogue, sound effects, and music synchronized with the picture in the same pass. In this mode the price is identical with audio on or off, so there's no cost reason to disable it.

Is Wan 3.0 reference-to-video free to try?+

You can try it free through Scenetra: new accounts get free welcome credits and a 7-day trial that cover your first generations. After that it's pay-per-generation with no subscription gate — a 5-second 480p test with your references costs $0.20.

Start creating with Wan 3.0

Use Wan 3.0 alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.

Get Started Free