Reference-to-Video Generation with Seedance 2.0
Generate videos from up to 9 reference images, 3 reference videos, and 3 audio tracks with Seedance 2.0 by ByteDance. Mention references in your prompt — 'image 1', 'video 1' — for character consistency, motion transfer, and audio-driven lip-sync, at resolutions up to 1080p.
What makes Seedance 2 Reference-to-Video different
Up to 15 reference assets in one generation
Combine up to 9 reference images, 3 reference videos (15 seconds combined), and 3 audio tracks in a single request. Characters, wardrobe, sets, motion, and sound can all come from your own material instead of being re-described in every prompt.
Character consistency across a series
Reuse the same reference images across generations and recurring characters keep their identity from clip to clip — the foundation for episodic content, brand mascots, and AI influencers that look the same every time. Mention each asset in the prompt by order ('image 1', 'image 2') so the model knows which role it plays.
Motion transfer and extension from reference videos
Feed in up to 3 short clips and the model can match their camera movement, continue their action, or restyle them — reference videos are built for editing- and extension-style workflows, with 15 seconds of combined footage per request.
Audio-driven generation with lip-sync
Attach a voice line or music track alongside at least one reference image or video, and the generated performance follows it — speech drives lip-sync, and pacing follows the sound. (Driving a generation from audio alone, with no visual reference, is a Seedance 2.5 feature.)
Playground
A timelapse of a flower blooming in a sunlit meadow, cinematic quality
Drop images or click to upload
Try Seedance 2 Reference-to-Video in Scenetra
Open PlaygroundParameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt | text | Optional text prompt. Mention references in order — 'image 1', 'video 1', 'audio 1'. | — |
| Reference Images | image upload | Up to 9 reference images for character, style, and scene consistency. | — |
| Reference Videos | video upload | Up to 3 reference videos (mp4/mov, 15 seconds combined) for motion transfer, editing, and extension. | — |
| Reference Audio | audio upload | Up to 3 audio files (wav/mp3) for audio-driven generation and lip-sync. Requires at least one reference image or video. | — |
| Duration | select | Video duration in seconds.4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 | 5 |
| Resolution | select | Output video resolution. 1080p is produced via super-resolution.480p · 720p · 1080p | 720p |
| Aspect Ratio | select | Video aspect ratio. 'Auto' follows the primary reference media.auto · 21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16 | auto |
| Generate Audio | select | Generate synchronized audio with the video.false · true | true |
Pricing
From $0.11 per second
Pay per generation — only pay for what you use.
480p
$0.11/s
per second
720p
$0.24/s
per second
1080p
$0.54/s
per second
What a video costs
| Duration | 480p | 720p | 1080p |
|---|---|---|---|
| 5 seconds | $0.55 | $1.22 | $2.75 |
| 10 seconds | $2.43 | $5.47 |
Use cases
Characters & series
- Consistent recurring characters
- Episodic story content
- Brand mascots
- AI influencer clips
Style & motion transfer
- Match a reference look
- Reuse camera movement
- Continue existing footage
- Scene-consistent shots
Music & audio-driven
- Lip-synced performances
- Beat-matched edits
- Voiceover-led scenes
- Sound-effect moments
Brand & marketing
- On-brand product videos
- Campaign variations
- Wardrobe/set continuity
- Multi-asset compositing
Related models
Seedance 2.5 Reference-to-Video
Generate videos from up to 30 reference images, 10 reference videos, and 10 audio tracks with Seedance 2.5 by ByteDance. Cite references directly in your prompt as @Image1, @Video1, @Audio1. Unique to 2.5: audio-only referencing — a single music or voice track can drive visual pacing, beat matching, and lip-sync.
View model →Seedance 2 Mini Reference-to-Video
Seedance 2 Mini Reference-to-Video generates consistent videos from up to 9 reference images, 3 reference videos, and 3 audio tracks. Reference subjects directly in your prompt as 'image 1', 'video 1', etc. for character consistency and motion transfer — at ByteDance's fast, affordable Mini pricing.
View model →Happy Horse Reference-to-Video
Alibaba's Happy Horse reference-to-video generates videos from 1-9 reference images. Reference each subject in your prompt as character1, character2, ... character9 (order matches your uploaded images) for consistent multi-character scenes. 720p or 1080p output, 3-15 second durations.
View model →Grok Reference-to-Video
Generate videos with consistent subject appearance from up to 7 reference images using xAI's Grok Imagine Video. Reference each image in your prompt as @Image1, @Image2, etc. for style and content guidance.
View model →Frequently asked questions
What is Seedance 2.0 Reference to Video?+
It's the multimodal mode of ByteDance's Seedance 2.0: instead of starting from just a prompt, you guide the generation with your own reference material — images for characters and style, videos for motion, audio for lip-sync and pacing. On Scenetra it runs as a node in the visual workflow editor, with output up to 1080p.
How many reference files can I use?+
Up to 9 reference images, up to 3 reference videos (mp4 or mov, 15 seconds combined), and up to 3 audio tracks (wav or mp3). At least one reference image or video is required per generation; audio references can't run alone.
How do I reference my uploads in the prompt?+
Mention them by type and submission order: 'image 1' is your first reference image, 'video 1' your first video, 'audio 1' your first audio track. For example: 'The character from image 1 walks through the scene in image 2, moving like video 1.' The prompt itself is optional, but naming the references tells the model which asset plays which role.
Can Seedance 2.0 extend or edit an existing video?+
Yes, within reference limits: supply your footage as a reference video (up to 3 clips, 15 seconds combined) and prompt a continuation, a restyle, or motion transfer onto new subjects. For longer combined reference footage — up to 30 seconds — Seedance 2.5 Reference to Video raises the caps.
How much does Reference to Video cost?+
The same per-second rate as the other Seedance 2.0 modes: about $0.11/s at 480p, $0.24/s at 720p, and $0.54/s at 1080p of generated video. A 5-second 720p generation is about $1.22; 10 seconds at 1080p about $5.47.
How do I keep a character consistent across multiple videos?+
Reuse the same reference images across generations. Upload your character from a few angles, mention them as 'image 1' through 'image 3' in every prompt, and the model keeps the identity stable from clip to clip — which is how episodic series and brand characters stay recognizable.
Start creating with Seedance 2 Reference-to-Video
Use Seedance 2 Reference-to-Video alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started Free