Text to VideoImage to VideoReference to Video

Reference-to-Video Generation with Seedance 2 Fast

Generate videos from your own reference material with Seedance 2 Fast by ByteDance — up to 9 reference images, 3 reference videos, and 3 audio tracks per generation. Mention references in your prompt as 'image 1', 'video 1', and so on for character consistency, motion transfer, and audio-driven clips at the fast tier's lower cost.

What makes Seedance 2 Fast Reference-to-Video different

Up to 15 reference assets in one generation

Combine up to 9 reference images, 3 reference videos (mp4 or mov, 15 seconds combined), and 3 audio tracks in a single request. Characters, wardrobe, sets, motion, and sound come from your own material instead of being re-described in every prompt — at the Fast tier's lower cost per second.

Character consistency for series work

Reuse the same reference images across generations and recurring characters keep their identity from clip to clip. Mention each asset in the prompt by order — 'image 1', 'image 2' — so the model knows which role it plays. Fast's cheaper rate makes it the natural tier for producing series volume.

Editing and extension from reference videos

Feed in up to 3 short clips and the model can continue their action, restyle them, or transfer their motion onto new subjects — reference videos are built for editing- and extension-style workflows, with 15 seconds of combined footage per request. Audio references require at least one image or video reference alongside them.

Draft references on Fast, finish on Seedance 2

The full Seedance 2 Reference to Video offers the same 9-image/3-video/3-audio workflow with a 1080p output tier. On a Scenetra board the two run side by side: dial in your reference setup and prompt on Fast, then re-run the final take through Seedance 2 when the deliverable needs full HD.

Playground

Seedance 2 Fast Reference-to-Video
playground

A timelapse of a flower blooming in a sunlit meadow, cinematic quality

Drop images or click to upload

5
720p
Auto
true
Faces
Generate
Resultpreview
Sample output from Seedance 2 Fast Reference-to-Video

Try Seedance 2 Fast Reference-to-Video in Scenetra

Open Playground

Parameters

ParameterTypeDescriptionDefault
Prompt*textText prompt describing the video. Mention references in order — 'image 1', 'video 1', 'audio 1'.
Reference Imagesimage uploadUp to 9 reference images for character, style, and scene consistency.
Reference Videosvideo uploadUp to 3 reference videos (mp4/mov, 15 seconds combined) for editing and extension.
Reference Audioaudio uploadUp to 3 audio files (wav/mp3). Requires at least one reference image or video.
DurationselectDuration of the video in seconds (4-15).4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 155
ResolutionselectVideo resolution. Higher resolutions cost more per second.480p · 720p720p
Aspect RatioselectAspect ratio. 'Auto' uses the primary reference media's aspect ratio.Auto · 21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16Auto
Generate AudioselectWhether to generate synchronized audio — voice, sound effects, and music — with the video.false · truetrue
ContentselectWhat your video contains. 'No faces' can unlock lower prices; 'Mature' relaxes content restrictions.No faces · Faces · MatureFaces

Pricing

From $0.09 per second

Pay per generation — only pay for what you use.

480p

$0.09/s

per second

720p

$0.20/s

per second

What a video costs

Duration480p720p
5 seconds$0.45$0.98
10 seconds$0.90$1.94

Use cases

Characters & series

  • Consistent recurring characters
  • Episodic story content
  • Brand mascots
  • AI influencer clips

Style & motion transfer

  • Match a reference look
  • Reuse camera movement
  • Continue existing footage
  • Scene-consistent shots

Music & audio-driven

  • Voiceover-led scenes
  • Beat-matched edits
  • Lip-synced performances
  • Sound-effect moments

Brand & marketing

  • On-brand product videos
  • Campaign variations at volume
  • Wardrobe and set continuity
  • Multi-asset compositing

Related models

Frequently asked questions

What is Seedance 2 Fast Reference to Video?+

It's the multimodal mode of ByteDance's Seedance 2 Fast — the quick, lower-cost Seedance 2 tier. Instead of starting from just a prompt, you guide the generation with your own reference material: images for characters and style, videos for editing and extension, audio for sound-driven clips. On Scenetra it runs as a node in the visual workflow editor at 480p or 720p.

How many reference files can I use?+

Up to 9 reference images, up to 3 reference videos (mp4 or mov, 15 seconds combined), and up to 3 audio tracks (wav or mp3). Audio references require at least one reference image or video alongside them.

How do I reference my uploads in the prompt?+

Mention them by type and submission order: 'image 1' is your first reference image, 'video 1' your first video, 'audio 1' your first audio track. For example: 'The character from image 1 walks through the scene in image 2.' Naming the references tells the model which asset plays which role.

How much does Seedance 2 Fast Reference to Video cost?+

The same per-second rates as the other Fast modes: about $0.09 per second at 480p and $0.20 per second at 720p, audio included. A 5-second 720p generation is about $0.98; 10 seconds about $1.94. One thing to know: when you supply reference videos, their duration can count toward the billed total alongside the output — so a generation with 10 seconds of reference footage can cost more than one from images alone. Scenetra bills at provider cost with 0% markup, and the node shows the exact price for your settings before you run it.

Can Seedance 2 Fast extend or edit an existing video?+

Yes, within reference limits: supply your footage as a reference video (up to 3 clips, 15 seconds combined) and prompt a continuation, a restyle, or motion transfer onto new subjects. For higher-resolution results, the full Seedance 2 Reference to Video offers the same workflow with a 1080p tier.

How do I keep a character consistent across multiple videos?+

Reuse the same reference images across generations. Upload your character from a few angles, mention them as 'image 1' through 'image 3' in every prompt, and the model keeps the identity stable from clip to clip — and Fast's lower per-second rate makes generating a whole series of takes affordable.

Does Reference to Video generate audio too?+

Yes — like the other Fast modes, it generates synchronized voice, sound effects, and music with the picture by default, and you can toggle audio off for silent output. You can also steer the sound yourself by attaching reference audio tracks alongside at least one visual reference.

Start creating with Seedance 2 Fast Reference-to-Video

Use Seedance 2 Fast Reference-to-Video alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.

Get Started Free