Seedance 2 Reference to Video
Generate videos from up to 9 reference images, 3 reference videos, and 3 audio tracks with Seedance 2.0 by ByteDance. Mention references in your prompt — 'image 1', 'video 1' — for character consistency, motion transfer, and audio-driven lip-sync, at resolutions up to 1080p.
40% less than the priciest provider (5 compared). Same model, same output.
Up to 9 reference images for character, style or scene. Cite in the prompt as @Image1 and so on.
Up to 3 reference videos (mp4/mov), total duration up to 15s. Reference video seconds are billed as input.
Up to 3 reference audio files (wav/mp3). Requires at least one reference image or video.
Duration
Duration of the video in seconds (4-15). Billed per second.
Resolution
Video resolution. Higher resolutions cost more per second.
Aspect Ratio
Aspect ratio of the generated video. 'adaptive' follows the prompt or the primary media.
Audio
Synchronized audio: voice, sound effects and music generated in the same pass. Some sellers cannot switch audio off, so turning it off can change the winning price.
Content
What the video contains. Sellers differ in what they allow: 'general' (no faces) can unlock lower prices; 'mature' routes only to sellers without content restrictions.
40% less than the priciest provider (5 compared). Same model, same output.
Your result renders here.
Use Seedance 2 Reference to Video via the API
Generate with Seedance 2 Reference to Video from your own code. Async by design: submit a generation, get a run id back instantly, poll until it completes. Billed in credits at the best available provider price — from $0.11 per second, audio included.
1 — Start a generation
curl -X POST https://app.scenetra.com/api/v1/generate/seedance-2-ref2v \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"prompt":"@Image1 walks through a sunlit market holding @Image2, handheld follow shot, crowd murmur","reference_images":["https://example.com/person.jpg","https://example.com/product.jpg"],"duration":"5"}'Response (202)
{
"runId": "run_abc123",
"status": "PENDING",
"model": "seedance-2-ref2v",
"estimatedCredits": 900,
"statusUrl": "/api/v1/runs/run_abc123"
}2 — Poll until COMPLETED (every 2-5 seconds)
curl https://app.scenetra.com/api/v1/runs/run_abc123 \
-H "Authorization: Bearer YOUR_API_KEY"Response
{
"runId": "run_abc123",
"status": "COMPLETED",
"progress": { "completed": 1, "total": 1 },
"outputs": [
{ "type": "video", "url": "https://images.scenetra.com/..." }
],
"error": null,
"estimatedCredits": 900
}Request fields
| Field | Type | Description | Default |
|---|---|---|---|
| prompt* | string | Cite reference inputs in order with @-syntax: @Image1, @Video1, @Audio1. | — |
| reference_images | string[] (https URLs) | Up to 9 reference images for character, style or scene. Cite in the prompt as @Image1 and so on. | — |
| reference_videos | string[] (https URLs) | Up to 3 reference videos (mp4/mov), total duration up to 15s. Reference video seconds are billed as input. | — |
| reference_audios | string[] (https URLs) | Up to 3 reference audio files (wav/mp3). Requires at least one reference image or video. | — |
| duration | string | Duration of the video in seconds (4-15). Billed per second.4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 | 5 |
| resolution | string | Video resolution. Higher resolutions cost more per second.480p · 720p · 1080p · 4k | 720p |
| aspect_ratio | string | Aspect ratio of the generated video. 'adaptive' follows the prompt or the primary media.adaptive · 21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16 | adaptive |
| generate_audio | string | Synchronized audio: voice, sound effects and music generated in the same pass. Some sellers cannot switch audio off, so turning it off can change the winning price.false · true | true |
| subject_content | string | What the video contains. Sellers differ in what they allow: 'general' (no faces) can unlock lower prices; 'mature' routes only to sellers without content restrictions.general · people · mature | people |
| callbackUrl | string (https URL) | Optional. When the run finishes (completed or failed), we POST the final run status to this URL — same shape as the poll response, plus model, params, and chargedCredits. Best-effort; polling remains the source of truth. | |
Unknown fields are rejected with a helpful 400. Media fields take https URLs. There is no provider parameter — every run is routed to the cheapest capable provider automatically.
Good to know
Billing
Runs bill your account credits at provider cost with 0% markup. The submit response includes the credit estimate; failed generations are refunded automatically. Add ?dryRun=1 to get the estimate without running anything.
Errors
Errors are always { "error": { "code", "message" } } with stable codes: unauthorized, model_not_available, invalid_parameter, insufficient_credits, rate_limited. A failed run is not an HTTP error — poll responses report status: "FAILED" with the reason.
What makes Seedance 2 Reference to Video different
Synchronized audio, generated with the picture
Seedance 2.0 creates voice, sound effects, and background music together with the video in a single pass — no separate audio model, no lip-sync post-processing. Dialogue matches mouths and footsteps land on frames because sound and picture come from the same generation. Flip one toggle for silent footage.
A real 1080p tier
Alongside 480p and 720p, Seedance 2.0 offers full-HD 1080p output via super-resolution — something the newer Seedance 2.5 does not (it currently caps at 720p). When the deliverable has to be crisp at full screen, 2.0 is the Seedance generation to reach for. Draft cheap at 480p, then re-run the winning prompt at 1080p.
4 to 15 seconds, priced per second
Pick any whole-second duration from 4 to 15. Billing is per second of generated video, so a 4-second draft costs a fraction of a 15-second final — iterate short, commit long. Aspect ratios span 21:9 cinematic wide to 9:16 vertical, or leave it on auto and let the prompt decide.
One model, three modes on one board
Text to Video is one of three Seedance 2.0 modes on Scenetra. Animate a first-frame image (with optional last-frame control) using Image to Video, or drive generations from up to 9 reference images, 3 videos, and 3 audio tracks with Reference to Video. All three run as nodes you can wire into upscalers, sequencers, and other models on the same canvas.
More made with Seedance 2 Reference to Video
Parameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Text prompt describing the desired video. | — |
| Duration | select | Video duration in seconds.4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 | 5 |
| Resolution | select | Output video resolution. 1080p is produced via super-resolution.480p · 720p · 1080p | 720p |
| Aspect Ratio | select | Video aspect ratio, or 'auto' to let the model choose from prompt context.auto · 21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16 | auto |
| Generate Audio | select | Generate synchronized audio — voice, sound effects, and background music — with the video.false · true | true |
Pricing
From $0.11 per second, audio included
Pay per generation — only pay for what you use.
480p
$0.11/s
per second
720p
$0.24/s
per second
1080p
$0.54/s
per second
What a video costs
| Duration | 480p | 720p | 1080p |
|---|---|---|---|
| 5 seconds | $0.55 | $1.22 | $2.75 |
| 10 seconds | $2.43 | $5.47 |
Use cases
Film & storytelling
- Cinematic establishing shots
- Dialogue scenes with lip-sync
- Previz and animatics
- Short narrative beats
Marketing & ads
- Product spots with music
- UGC-style ad concepts
- Brand stings in 21:9
- A/B concept testing at 480p
Social content
- Vertical 9:16 clips
- Sound-on hooks
- Trend-reactive videos
- Series with a consistent look
Product & e-commerce
- Product showcase clips
- Lifestyle b-roll
- Seasonal campaign variants
- 1080p hero videos
Frequently asked questions
What is Seedance 2.0 Reference to Video?+
It's the multimodal mode of ByteDance's Seedance 2.0: instead of starting from just a prompt, you guide the generation with your own reference material — images for characters and style, videos for motion, audio for lip-sync and pacing. On Scenetra it runs as a node in the visual workflow editor, with output up to 1080p.
How many reference files can I use?+
Up to 9 reference images, up to 3 reference videos (mp4 or mov, 15 seconds combined), and up to 3 audio tracks (wav or mp3). At least one reference image or video is required per generation; audio references can't run alone.
How do I reference my uploads in the prompt?+
Mention them by type and submission order: 'image 1' is your first reference image, 'video 1' your first video, 'audio 1' your first audio track. For example: 'The character from image 1 walks through the scene in image 2, moving like video 1.' The prompt itself is optional, but naming the references tells the model which asset plays which role.
Can Seedance 2.0 extend or edit an existing video?+
Yes, within reference limits: supply your footage as a reference video (up to 3 clips, 15 seconds combined) and prompt a continuation, a restyle, or motion transfer onto new subjects. For longer combined reference footage — up to 30 seconds — Seedance 2.5 Reference to Video raises the caps.
How do I keep a character consistent across multiple videos?+
Reuse the same reference images across generations. Upload your character from a few angles, mention them as 'image 1' through 'image 3' in every prompt, and the model keeps the identity stable from clip to clip — which is how episodic series and brand characters stay recognizable.
One-click tools built on Seedance 2 Reference to Video
Ready-made workflows — fill in a form, no node editor needed.
AI Product UGC Videos
Put your product in real-looking hands — upload product photos and up to four model photos, and get four UGC-style testimonial videos of models holding and presenting it.
Open tool →AI UGC Video Generator
Turn one face photo into four ready-to-post UGC-style videos — street interview, living-room dance, kitchen talk, and garden chat — starring a consistent AI creator.
Open tool →Seedance 2 Video Generator
Generate cinematic AI videos with synchronized audio from a text prompt, powered by ByteDance's Seedance 2.0 — up to 15 seconds, 1080p, any aspect ratio.
Open tool →Related models
Seedance 2.5
ByteDance's Seedance 2.5 generates cinematic videos up to 30 seconds from text prompts with synchronized audio — voice, sound effects, and background music — built in by default. The newest Seedance generation, with mp4 or high-fidelity mov output for multi-round editing pipelines.
View model →Seedance 2 Mini
ByteDance's Seedance 2 Mini is the fast, lower-cost tier of Seedance 2 — generating cinematic videos from text prompts with synchronized audio. Mini delivers the same audio-visual sync at a fraction of the price, ideal for rapid iteration and high-volume social content.
View model →Kling 3.0 Pro
Kuaishou's flagship video generation model delivering stunning visual fidelity. Kling 3.0 Pro supports text-to-video, image-to-video with start/end frames, element references for character consistency, and native audio generation.
View model →Start creating with Seedance 2 Reference to Video
Use Seedance 2 Reference to Video alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started