Text to VideoImage to VideoReference to Video

Reference-to-Video Generation with Seedance 2.0

Generate videos from up to 9 reference images, 3 reference videos, and 3 audio tracks with Seedance 2.0 by ByteDance. Mention references in your prompt — 'image 1', 'video 1' — for character consistency, motion transfer, and audio-driven lip-sync, at resolutions up to 1080p.

What makes Seedance 2 Reference-to-Video different

Up to 15 reference assets in one generation

Combine up to 9 reference images, 3 reference videos (15 seconds combined), and 3 audio tracks in a single request. Characters, wardrobe, sets, motion, and sound can all come from your own material instead of being re-described in every prompt.

Character consistency across a series

Reuse the same reference images across generations and recurring characters keep their identity from clip to clip — the foundation for episodic content, brand mascots, and AI influencers that look the same every time. Mention each asset in the prompt by order ('image 1', 'image 2') so the model knows which role it plays.

Motion transfer and extension from reference videos

Feed in up to 3 short clips and the model can match their camera movement, continue their action, or restyle them — reference videos are built for editing- and extension-style workflows, with 15 seconds of combined footage per request.

Audio-driven generation with lip-sync

Attach a voice line or music track alongside at least one reference image or video, and the generated performance follows it — speech drives lip-sync, and pacing follows the sound. (Driving a generation from audio alone, with no visual reference, is a Seedance 2.5 feature.)

Playground

Seedance 2 Reference-to-Video
playground

A timelapse of a flower blooming in a sunlit meadow, cinematic quality

Drop images or click to upload

5
720p
auto
true
Generate
Resultpreview
Sample output from Seedance 2 Reference-to-Video

Try Seedance 2 Reference-to-Video in Scenetra

Open Playground

Parameters

ParameterTypeDescriptionDefault
PrompttextOptional text prompt. Mention references in order — 'image 1', 'video 1', 'audio 1'.
Reference Imagesimage uploadUp to 9 reference images for character, style, and scene consistency.
Reference Videosvideo uploadUp to 3 reference videos (mp4/mov, 15 seconds combined) for motion transfer, editing, and extension.
Reference Audioaudio uploadUp to 3 audio files (wav/mp3) for audio-driven generation and lip-sync. Requires at least one reference image or video.
DurationselectVideo duration in seconds.4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 155
ResolutionselectOutput video resolution. 1080p is produced via super-resolution.480p · 720p · 1080p720p
Aspect RatioselectVideo aspect ratio. 'Auto' follows the primary reference media.auto · 21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16auto
Generate AudioselectGenerate synchronized audio with the video.false · truetrue

Pricing

From $0.11 per second

Pay per generation — only pay for what you use.

480p

$0.11/s

per second

720p

$0.24/s

per second

1080p

$0.54/s

per second

What a video costs

Duration480p720p1080p
5 seconds$0.55$1.22$2.75
10 seconds$2.43$5.47

Use cases

Characters & series

  • Consistent recurring characters
  • Episodic story content
  • Brand mascots
  • AI influencer clips

Style & motion transfer

  • Match a reference look
  • Reuse camera movement
  • Continue existing footage
  • Scene-consistent shots

Music & audio-driven

  • Lip-synced performances
  • Beat-matched edits
  • Voiceover-led scenes
  • Sound-effect moments

Brand & marketing

  • On-brand product videos
  • Campaign variations
  • Wardrobe/set continuity
  • Multi-asset compositing

Related models

Frequently asked questions

What is Seedance 2.0 Reference to Video?+

It's the multimodal mode of ByteDance's Seedance 2.0: instead of starting from just a prompt, you guide the generation with your own reference material — images for characters and style, videos for motion, audio for lip-sync and pacing. On Scenetra it runs as a node in the visual workflow editor, with output up to 1080p.

How many reference files can I use?+

Up to 9 reference images, up to 3 reference videos (mp4 or mov, 15 seconds combined), and up to 3 audio tracks (wav or mp3). At least one reference image or video is required per generation; audio references can't run alone.

How do I reference my uploads in the prompt?+

Mention them by type and submission order: 'image 1' is your first reference image, 'video 1' your first video, 'audio 1' your first audio track. For example: 'The character from image 1 walks through the scene in image 2, moving like video 1.' The prompt itself is optional, but naming the references tells the model which asset plays which role.

Can Seedance 2.0 extend or edit an existing video?+

Yes, within reference limits: supply your footage as a reference video (up to 3 clips, 15 seconds combined) and prompt a continuation, a restyle, or motion transfer onto new subjects. For longer combined reference footage — up to 30 seconds — Seedance 2.5 Reference to Video raises the caps.

How much does Reference to Video cost?+

The same per-second rate as the other Seedance 2.0 modes: about $0.11/s at 480p, $0.24/s at 720p, and $0.54/s at 1080p of generated video. A 5-second 720p generation is about $1.22; 10 seconds at 1080p about $5.47.

How do I keep a character consistent across multiple videos?+

Reuse the same reference images across generations. Upload your character from a few angles, mention them as 'image 1' through 'image 3' in every prompt, and the model keeps the identity stable from clip to clip — which is how episodic series and brand characters stay recognizable.

Start creating with Seedance 2 Reference-to-Video

Use Seedance 2 Reference-to-Video alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.

Get Started Free