Reference-to-Video with up to 30 References — Seedance 2.5
Generate videos from up to 30 reference images, 10 reference videos, and 10 audio tracks with Seedance 2.5 by ByteDance. Cite references directly in your prompt as @Image1, @Video1, @Audio1. Unique to 2.5: audio-only referencing — a single music or voice track can drive visual pacing, beat matching, and lip-sync.
What makes Seedance 2.5 Reference-to-Video different
Up to 50 reference assets in one generation
Combine up to 30 reference images, 10 reference videos (30 seconds total), and 10 audio tracks (30 seconds total) in a single request. Characters, wardrobe, sets, motion, and sound can all come from your own material — the most reference capacity of any Seedance generation.
Audio-only referencing — unique to 2.5
Drop in a single music track, voice line, or sound effect with no visuals at all, and the model builds the video around it: visual pacing follows the beat, cuts land on transitions, and speech drives lip-sync. No other Seedance mode can be driven by sound alone.
Cite references right in the prompt
Reference inputs by submission order with @-syntax — 'The character in @Image1 dances to @Audio1 in the courtyard from @Image2'. The model knows exactly which asset plays which role, instead of guessing from a pile of attachments.
Character consistency across a series
Because the same reference images can be reused across generations, recurring characters keep their identity from clip to clip — the foundation for episodic content, brand mascots, and AI influencers that look the same every time.
Playground
Try Seedance 2.5 Reference-to-Video right here
Fill in the inputs and hit Generate — your settings carry straight into Scenetra, where the video renders in about two minutes. New accounts start with free credits.
Characters, style, or scenes — cite them in the prompt as @Image1, @Image2…
Describe the video, citing references with @-syntax (@Image1, @Audio1).
Optional — a music or voice track can drive pacing and lip-sync.
Seconds of video (4-30).
Parameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt | text | Cite reference inputs in submission order with @-syntax: @Image1, @Video1, @Audio1. | — |
| Reference Images | image upload | Up to 30 reference images for character, style, and scene consistency. | — |
| Reference Videos | video upload | Up to 10 reference videos (combined duration up to 30s) for motion and style transfer. | — |
| Reference Audio | audio upload | Up to 10 audio files (combined up to 30s). Audio-only referencing can drive pacing, beats, and lip-sync. | — |
| Duration | select | Video duration from 4 to 30 seconds. | 5 |
| Resolution | select | Output video resolution.480p · 720p | 720p |
| Aspect Ratio | select | Video aspect ratio; 'adaptive' follows the primary reference media.16:9 · 4:3 · 1:1 · 3:4 · 9:16 · 21:9 · adaptive | adaptive |
| Generate Audio | select | Synchronized audio generated with the video.false · true | true |
Pricing
From $0.14 per second, audio included
Pay per generation — only pay for what you use.
480p
$0.14/s
per second
720p
$0.30/s
per second
What a video costs
| Duration | 480p | 720p |
|---|---|---|
| 5 seconds | $0.70 | $1.51 |
| 10 seconds | $1.40 | $3.02 |
| 30 seconds | $4.20 | $9.03 |
Use cases
Characters & series
- Consistent recurring characters
- Episodic story content
- Brand mascots
- AI influencer clips
Music & audio-driven
- Beat-matched edits
- Lip-synced performances
- Lyric visualizers
- Sound-effect scenes
Style & motion transfer
- Match a reference look
- Reuse camera movement
- Continue existing footage
- Scene-consistent shots
Brand & marketing
- On-brand product videos
- Campaign variations
- Wardrobe/set continuity
- Multi-asset compositing
Related models
Seedance 2 Reference-to-Video
Generate videos from up to 9 reference images, 3 reference videos, and 3 audio tracks with Seedance 2.0 by ByteDance. Mention references in your prompt — 'image 1', 'video 1' — for character consistency, motion transfer, and audio-driven lip-sync, at resolutions up to 1080p.
View model →Seedance 2 Mini Reference-to-Video
Seedance 2 Mini Reference-to-Video generates consistent videos from up to 9 reference images, 3 reference videos, and 3 audio tracks. Reference subjects directly in your prompt as 'image 1', 'video 1', etc. for character consistency and motion transfer — at ByteDance's fast, affordable Mini pricing.
View model →Happy Horse Reference-to-Video
Alibaba's Happy Horse reference-to-video generates videos from 1-9 reference images. Reference each subject in your prompt as character1, character2, ... character9 (order matches your uploaded images) for consistent multi-character scenes. 720p or 1080p output, 3-15 second durations.
View model →Veo 3.1
Google's premier video generation model supporting text-to-video, image-to-video, and first-last-frame video creation. Veo 3.1 delivers up to 4K resolution with native audio synthesis, making it one of the most versatile video models available.
View model →Frequently asked questions
What is Seedance 2.5 Reference to Video?+
It's the multimodal mode of ByteDance's Seedance 2.5: instead of starting from just a prompt, you guide the generation with your own reference material — images for characters and style, videos for motion, audio for pacing and lip-sync. Cite each asset in the prompt as @Image1, @Video1, @Audio1. On Scenetra it runs as a node in the visual workflow editor.
How many reference files can I use?+
Up to 30 reference images, up to 10 reference videos with a combined duration of 30 seconds, and up to 10 audio tracks, also 30 seconds combined — up to 50 assets per generation. Each image can be up to 30MB (300-6000px per side); videos are mp4/mov clips of 2-30 seconds each; audio is wav or mp3.
What is audio-only referencing?+
A Seedance 2.5 exclusive: you can drive a generation with sound alone. Supply one music track, voice recording, or sound effect — no reference images required — and the model generates visuals that follow it: pacing matches the beat, and speech produces lip-synced performance. It's the fastest way to turn a song or voiceover into video.
How do I reference my uploads in the prompt?+
Use @-syntax in submission order: @Image1 is your first image, @Video1 your first video, @Audio1 your first audio track. For example: 'The character from @Image1 walks through the city in @Image2, moving like @Video1, scored by @Audio1.' The model maps each mention to the exact asset.
Can Seedance 2.5 extend or edit an existing video?+
You can supply existing footage as a reference video and prompt a continuation or a stylistic match — that covers most extension-style use cases. The model's dedicated frame-level video-editing mode is not yet exposed on Scenetra.
How much does Reference to Video cost?+
The base rate matches the other Seedance 2.5 modes: about $0.14/s at 480p and $0.30/s at 720p of generated video. One thing to know: when you attach reference videos, their footage is processed too, so generations with long video references cost more than the same duration without them. Reference images and audio don't add duration-based cost.
How do I keep a character consistent across multiple videos?+
Reuse the same reference images across generations. Upload your character from a few angles, cite them as @Image1-@Image3 in every prompt, and the model keeps the identity stable from clip to clip — which is how episodic series and brand characters stay recognizable.
Can I use photos of real people as references?+
Directly uploading images or videos containing real human faces is not permitted by the model. Use model-generated characters, preset digital characters, or properly authorized real-person assets instead.
Start creating with Seedance 2.5 Reference-to-Video
Use Seedance 2.5 Reference-to-Video alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started Free