Reference-to-Video Generation with Seedance 2 Fast
Generate videos from your own reference material with Seedance 2 Fast by ByteDance — up to 9 reference images, 3 reference videos, and 3 audio tracks per generation. Mention references in your prompt as 'image 1', 'video 1', and so on for character consistency, motion transfer, and audio-driven clips at the fast tier's lower cost.
What makes Seedance 2 Fast Reference-to-Video different
Up to 15 reference assets in one generation
Combine up to 9 reference images, 3 reference videos (mp4 or mov, 15 seconds combined), and 3 audio tracks in a single request. Characters, wardrobe, sets, motion, and sound come from your own material instead of being re-described in every prompt — at the Fast tier's lower cost per second.
Character consistency for series work
Reuse the same reference images across generations and recurring characters keep their identity from clip to clip. Mention each asset in the prompt by order — 'image 1', 'image 2' — so the model knows which role it plays. Fast's cheaper rate makes it the natural tier for producing series volume.
Editing and extension from reference videos
Feed in up to 3 short clips and the model can continue their action, restyle them, or transfer their motion onto new subjects — reference videos are built for editing- and extension-style workflows, with 15 seconds of combined footage per request. Audio references require at least one image or video reference alongside them.
Draft references on Fast, finish on Seedance 2
The full Seedance 2 Reference to Video offers the same 9-image/3-video/3-audio workflow with a 1080p output tier. On a Scenetra board the two run side by side: dial in your reference setup and prompt on Fast, then re-run the final take through Seedance 2 when the deliverable needs full HD.
Playground
A timelapse of a flower blooming in a sunlit meadow, cinematic quality
Drop images or click to upload
Try Seedance 2 Fast Reference-to-Video in Scenetra
Open PlaygroundParameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Text prompt describing the video. Mention references in order — 'image 1', 'video 1', 'audio 1'. | — |
| Reference Images | image upload | Up to 9 reference images for character, style, and scene consistency. | — |
| Reference Videos | video upload | Up to 3 reference videos (mp4/mov, 15 seconds combined) for editing and extension. | — |
| Reference Audio | audio upload | Up to 3 audio files (wav/mp3). Requires at least one reference image or video. | — |
| Duration | select | Duration of the video in seconds (4-15).4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 | 5 |
| Resolution | select | Video resolution. Higher resolutions cost more per second.480p · 720p | 720p |
| Aspect Ratio | select | Aspect ratio. 'Auto' uses the primary reference media's aspect ratio.Auto · 21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16 | Auto |
| Generate Audio | select | Whether to generate synchronized audio — voice, sound effects, and music — with the video.false · true | true |
| Content | select | What your video contains. 'No faces' can unlock lower prices; 'Mature' relaxes content restrictions.No faces · Faces · Mature | Faces |
Pricing
From $0.09 per second
Pay per generation — only pay for what you use.
480p
$0.09/s
per second
720p
$0.20/s
per second
What a video costs
| Duration | 480p | 720p |
|---|---|---|
| 5 seconds | $0.45 | $0.98 |
| 10 seconds | $0.90 | $1.94 |
Use cases
Characters & series
- Consistent recurring characters
- Episodic story content
- Brand mascots
- AI influencer clips
Style & motion transfer
- Match a reference look
- Reuse camera movement
- Continue existing footage
- Scene-consistent shots
Music & audio-driven
- Voiceover-led scenes
- Beat-matched edits
- Lip-synced performances
- Sound-effect moments
Brand & marketing
- On-brand product videos
- Campaign variations at volume
- Wardrobe and set continuity
- Multi-asset compositing
Related models
Seedance 2 Reference-to-Video
Generate videos from up to 9 reference images, 3 reference videos, and 3 audio tracks with Seedance 2.0 by ByteDance. Mention references in your prompt — 'image 1', 'video 1' — for character consistency, motion transfer, and audio-driven lip-sync, at resolutions up to 1080p.
View model →Seedance 2 Mini Reference-to-Video
Seedance 2 Mini Reference-to-Video generates consistent videos from up to 9 reference images, 3 reference videos, and 3 audio tracks. Reference subjects directly in your prompt as 'image 1', 'video 1', etc. for character consistency and motion transfer — at ByteDance's fast, affordable Mini pricing.
View model →Seedance 2.5 Reference-to-Video
Generate videos from up to 30 reference images, 10 reference videos, and 10 audio tracks with Seedance 2.5 by ByteDance. Cite references directly in your prompt as @Image1, @Video1, @Audio1. Unique to 2.5: audio-only referencing — a single music or voice track can drive visual pacing, beat matching, and lip-sync.
View model →Frequently asked questions
What is Seedance 2 Fast Reference to Video?+
It's the multimodal mode of ByteDance's Seedance 2 Fast — the quick, lower-cost Seedance 2 tier. Instead of starting from just a prompt, you guide the generation with your own reference material: images for characters and style, videos for editing and extension, audio for sound-driven clips. On Scenetra it runs as a node in the visual workflow editor at 480p or 720p.
How many reference files can I use?+
Up to 9 reference images, up to 3 reference videos (mp4 or mov, 15 seconds combined), and up to 3 audio tracks (wav or mp3). Audio references require at least one reference image or video alongside them.
How do I reference my uploads in the prompt?+
Mention them by type and submission order: 'image 1' is your first reference image, 'video 1' your first video, 'audio 1' your first audio track. For example: 'The character from image 1 walks through the scene in image 2.' Naming the references tells the model which asset plays which role.
How much does Seedance 2 Fast Reference to Video cost?+
The same per-second rates as the other Fast modes: about $0.09 per second at 480p and $0.20 per second at 720p, audio included. A 5-second 720p generation is about $0.98; 10 seconds about $1.94. One thing to know: when you supply reference videos, their duration can count toward the billed total alongside the output — so a generation with 10 seconds of reference footage can cost more than one from images alone. Scenetra bills at provider cost with 0% markup, and the node shows the exact price for your settings before you run it.
Can Seedance 2 Fast extend or edit an existing video?+
Yes, within reference limits: supply your footage as a reference video (up to 3 clips, 15 seconds combined) and prompt a continuation, a restyle, or motion transfer onto new subjects. For higher-resolution results, the full Seedance 2 Reference to Video offers the same workflow with a 1080p tier.
How do I keep a character consistent across multiple videos?+
Reuse the same reference images across generations. Upload your character from a few angles, mention them as 'image 1' through 'image 3' in every prompt, and the model keeps the identity stable from clip to clip — and Fast's lower per-second rate makes generating a whole series of takes affordable.
Does Reference to Video generate audio too?+
Yes — like the other Fast modes, it generates synchronized voice, sound effects, and music with the picture by default, and you can toggle audio off for silent output. You can also steer the sound yourself by attaching reference audio tracks alongside at least one visual reference.
Start creating with Seedance 2 Fast Reference-to-Video
Use Seedance 2 Fast Reference-to-Video alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started Free