Reference-to-Video with Alibaba Wan 3.0
Generate video from up to 10 reference images with Alibaba's Wan 3.0. The model carries the people, objects, and styles from your references into a continuous take of 2 to 30 seconds, with native synchronized audio generated in the same pass. Output at 480p, 720p, or 1080p — reference images are free, and audio adds nothing to the price.
What makes Wan 3.0 different
Up to 10 references, one consistent cast
Feed Wan 3.0 between 1 and 10 reference images and it carries those people, objects, and styles into the generated video. That's the difference between an image-to-video starting frame and true reference conditioning: the references don't have to appear as frame one — they define who and what shows up anywhere in the shot, so a character can walk into a scene that none of your images contain.
30-second takes with native audio
Like the other Wan 3.0 modes, reference-to-video generates any duration from 2 to 30 seconds as one continuous take, with dialogue, sound effects, and music synchronized in the same pass. In this mode audio truly costs nothing extra — the per-second rate is identical whether the toggle is on or off — and the reference images are free at every resolution.
Flat, predictable pricing
Reference-to-video is billed at a flat $0.04 per second at 480p, $0.08 at 720p, and $0.16 at 1080p — higher than the cheapest text- and image-to-video rates, which is the premium for reference conditioning. A 10-second 720p clip with a consistent character is $0.80; a full 30-second 1080p take is $4.80. Draft at 480p to lock the likeness before rendering finals.
Build a character pipeline on one board
On a Scenetra board, generate a character sheet with an image model, wire the outputs into the Wan 3.0 reference node, and produce shot after shot with the same face, wardrobe, and style — a repeatable character pipeline on one canvas. Everything is billed at provider cost with 0% markup, and new accounts get free welcome credits plus a 7-day trial.
Playground
A timelapse of a flower blooming in a sunlit meadow, cinematic quality
Try Wan 3.0 in Scenetra
Open PlaygroundParameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Describe the video: subject, motion, camera work, and audio direction. Native synchronized audio is generated in the same pass. | — |
| Reference Images* | image upload | 1-10 reference images contributing people, objects, or styles to the generated video. | — |
| Resolution | select | Output resolution. 480p is the budget rung; 1080p is the premium rung.480p · 720p · 1080p | 720p |
| Aspect Ratio | select | Aspect ratio of the generated video. Auto lets the model choose the best framing for your prompt.Auto · 16:9 · 4:3 · 1:1 · 3:4 · 9:16 | Auto |
| Duration | select | Length of the generated video in seconds. Wan 3.0 generates any length from 2 to 30 seconds in a single take.2 · 3 · 4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 · 16 · 17 · 18 · 19 · 20 · 21 · 22 · 23 · 24 · 25 · 26 · 27 · 28 · 29 · 30 | 5 |
| Audio | boolean | Generate synchronized audio (dialogue, SFX, music) in the same pass. Same price either way in this mode. | true |
Pricing
From $0.04 per second
Pay per generation — only pay for what you use.
480p
$0.04/s
per second
720p
$0.08/s
per second
1080p
$0.16/s
per second
Audio
Included
per second
Reference images
Free
per second
What a video costs
| Duration | 480p | 720p | 1080p |
|---|---|---|---|
| 5 seconds | $0.20 | $0.40 | $0.80 |
| 10 seconds | $0.40 | $0.80 | $1.60 |
Use cases
Character content
- Recurring characters across episodes
- Consistent faces shot after shot
- Virtual influencer clips
- Character dialogue scenes with audio
Brand & product
- Same product in new scenes
- Brand mascots in motion
- Style-matched campaign series
- Wardrobe and prop continuity
Film & previz
- Cast continuity in previz
- Scene variations with one look
- Style-frame-driven sequences
- 30-second scene blocking
Social series
- Serialized character shorts
- Consistent visual identity per feed
- Multi-part stories, one cast
- Sound-on episodes up to 30 seconds
Related models
Seedance 2.5 Reference-to-Video
Generate videos from up to 30 reference images, 10 reference videos, and 10 audio tracks with Seedance 2.5 by ByteDance. Cite references directly in your prompt as @Image1, @Video1, @Audio1. Unique to 2.5: audio-only referencing — a single music or voice track can drive visual pacing, beat matching, and lip-sync.
View model →MiniMax H3 Reference-to-Video
Direct MiniMax H3 (Hailuo 03) with the material you already have: up to 9 reference images, 3 reference video clips, and 3 audio tracks in a single generation. Keep subjects and products consistent across shots, borrow motion and pacing from reference footage, sync action to a music track, and get native audio in the output — 5-15 seconds, from 480P up to 4K.
View model →PixVerse v6
Drive PixVerse v6 with up to 7 reference images and bind them directly in your prompt as @image1, @image2. Keep characters, objects and scenes consistent across 1-15 second videos at 360p to 1080p, with eight aspect ratios and optional synchronized audio. From $0.0225 per second.
View model →Frequently asked questions
What is Wan 3.0 reference-to-video?+
It's the mode of Alibaba's Wan 3.0 that generates video from up to 10 reference images. Instead of animating one starting frame, the model uses your references to keep specific people, objects, and styles consistent throughout the generated clip — 2 to 30 seconds in a single take, with native synchronized audio. On Scenetra it runs as a node in the visual workflow editor.
How much does Wan 3.0 reference-to-video cost?+
$0.04 per second at 480p, $0.08 at 720p, and $0.16 at 1080p — flat rates, with the reference images free and audio included at no extra cost. A 5-second 720p clip is $0.40 and a 10-second one $0.80. Everything is billed at provider cost with 0% markup.
Why does reference-to-video cost more than Wan 3.0 text-to-video?+
Reference conditioning is billed on a different, flat price ladder. Text- and image-to-video with audio route to a cheaper rate (from $0.025 per second at 480p), while reference-to-video is $0.04 to $0.16 per second depending on resolution regardless of the audio setting. The premium buys consistency: the same faces, objects, and styles across every generation.
How many reference images can I use?+
Between 1 and 10 per generation, and they're free — you pay only the per-second video rate. You can mix reference types in one job: a character, a product, and a style frame together, with the prompt describing how they interact.
What's the difference between reference-to-video and image-to-video?+
Image-to-video animates one upload as the literal first frame of the clip. Reference-to-video accepts up to 10 images that condition the content — the people, objects, and styles they contain appear consistently anywhere in the shot, in scenes your images never showed. Use image-to-video to bring a finished composition to life, and reference-to-video to keep a cast or look consistent across many shots.
Does reference-to-video generate audio too?+
Yes — like the other Wan 3.0 modes it generates dialogue, sound effects, and music synchronized with the picture in the same pass. In this mode the price is identical with audio on or off, so there's no cost reason to disable it.
Is Wan 3.0 reference-to-video free to try?+
You can try it free through Scenetra: new accounts get free welcome credits and a 7-day trial that cover your first generations. After that it's pay-per-generation with no subscription gate — a 5-second 480p test with your references costs $0.20.
Start creating with Wan 3.0
Use Wan 3.0 alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started Free