Seedance 2.5 Reference to Video
Generate videos from up to 30 reference images, 10 reference videos, and 10 audio tracks with Seedance 2.5 by ByteDance. Cite references directly in your prompt as @Image1, @Video1, @Audio1. Unique to 2.5: audio-only referencing — a single music or voice track can drive visual pacing, beat matching, and lip-sync.
42% less than the priciest provider (5 compared). Same model, same output.
Up to 30 reference images for character, style or scene. Cite in the prompt as @Image1 and so on.
Up to 10 reference videos (mp4/mov), combined duration up to 30s. Cite as @Video1. Reference video seconds are billed as input.
Up to 10 reference audio files (wav/mp3), combined up to 30s. One track can drive pacing, beats and lip-sync.
Duration
Duration of the video in seconds (4-30). Billed per second.
Resolution
Video resolution.
Aspect Ratio
Aspect ratio of the generated video. 'adaptive' follows the prompt or the primary media.
Audio
Synchronized audio: voice, sound effects and music generated in the same pass.
Content
What the video contains. Sellers differ in what they allow: 'general' (no faces) can unlock lower prices; 'mature' routes only to sellers without content restrictions.
42% less than the priciest provider (5 compared). Same model, same output.
Your result renders here.
Use Seedance 2.5 Reference to Video via the API
Generate with Seedance 2.5 Reference to Video from your own code. Async by design: submit a generation, get a run id back instantly, poll until it completes. Billed in credits at the best available provider price — from $0.14 per second, audio included.
1 — Start a generation
curl -X POST https://app.scenetra.com/api/v1/generate/seedance-2-5-ref2v \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"prompt":"@Image1 walks through a sunlit market holding @Image2, handheld follow shot, crowd murmur","reference_images":["https://example.com/person.jpg","https://example.com/product.jpg"],"duration":"5"}'Response (202)
{
"runId": "run_abc123",
"status": "PENDING",
"model": "seedance-2-5-ref2v",
"estimatedCredits": 1350,
"statusUrl": "/api/v1/runs/run_abc123"
}2 — Poll until COMPLETED (every 2-5 seconds)
curl https://app.scenetra.com/api/v1/runs/run_abc123 \
-H "Authorization: Bearer YOUR_API_KEY"Response
{
"runId": "run_abc123",
"status": "COMPLETED",
"progress": { "completed": 1, "total": 1 },
"outputs": [
{ "type": "video", "url": "https://images.scenetra.com/..." }
],
"error": null,
"estimatedCredits": 1350
}Request fields
| Field | Type | Description | Default |
|---|---|---|---|
| prompt* | string | Cite reference inputs in order with @-syntax: @Image1, @Video1, @Audio1. | — |
| reference_images | string[] (https URLs) | Up to 30 reference images for character, style or scene. Cite in the prompt as @Image1 and so on. | — |
| reference_videos | string[] (https URLs) | Up to 10 reference videos (mp4/mov), combined duration up to 30s. Cite as @Video1. Reference video seconds are billed as input. | — |
| reference_audios | string[] (https URLs) | Up to 10 reference audio files (wav/mp3), combined up to 30s. One track can drive pacing, beats and lip-sync. | — |
| duration | string | Duration of the video in seconds (4-30). Billed per second.4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 · 16 · 17 · 18 · 19 · 20 · 21 · 22 · 23 · 24 · 25 · 26 · 27 · 28 · 29 · 30 | 5 |
| resolution | string | Video resolution.480p · 720p | 720p |
| aspect_ratio | string | Aspect ratio of the generated video. 'adaptive' follows the prompt or the primary media.adaptive · 21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16 | adaptive |
| generate_audio | string | Synchronized audio: voice, sound effects and music generated in the same pass.false · true | true |
| subject_content | string | What the video contains. Sellers differ in what they allow: 'general' (no faces) can unlock lower prices; 'mature' routes only to sellers without content restrictions.general · people · mature | people |
| callbackUrl | string (https URL) | Optional. When the run finishes (completed or failed), we POST the final run status to this URL — same shape as the poll response, plus model, params, and chargedCredits. Best-effort; polling remains the source of truth. | |
Unknown fields are rejected with a helpful 400. Media fields take https URLs. There is no provider parameter — every run is routed to the cheapest capable provider automatically.
Good to know
Billing
Runs bill your account credits at provider cost with 0% markup. The submit response includes the credit estimate; failed generations are refunded automatically. Add ?dryRun=1 to get the estimate without running anything.
Errors
Errors are always { "error": { "code", "message" } } with stable codes: unauthorized, model_not_available, invalid_parameter, insufficient_credits, rate_limited. A failed run is not an HTTP error — poll responses report status: "FAILED" with the reason.
What makes Seedance 2.5 Reference to Video different
Up to 30 seconds in a single generation
Most AI video models cap out at 5-15 seconds. Seedance 2.5 generates up to 30 seconds in one pass — enough for a complete scene with a beginning, middle, and end. No stitching, no continuity drift between clips, no re-prompting to extend. Pick any whole-second duration from 4 to 30.
Synchronized audio, on by default
Every generation ships with native audio: spoken dialogue, ambient sound effects, and background music that match what's on screen. There is no separate audio pass and no lip-sync post-processing — the model generates picture and sound together. Turn it off with a single toggle if you want silent footage.
Built for editing pipelines
Alongside standard mp4, Seedance 2.5 can output mov encoded as yuv444p — higher color fidelity that holds up through multi-round editing and extension workflows where recompression loss normally accumulates. Feed the output back into an edit, upscale, or sequence node on your Scenetra board without visible degradation.
One model, three modes
Text to Video is one of three Seedance 2.5 modes on Scenetra. Start from a first-frame image with Image to Video, or drive generations from up to 30 reference images, 10 videos, and 10 audio tracks with Reference to Video — including audio-only referencing, where a single music track drives pacing and lip-sync.
More made with Seedance 2.5 Reference to Video
Parameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Text prompt describing the desired video. Supports English and Chinese. | — |
| Duration | select | Video duration from 4 to 30 seconds. | 5 |
| Resolution | select | Output video resolution.480p · 720p | 720p |
| Aspect Ratio | select | Video aspect ratio, or 'adaptive' to select automatically from prompt context.16:9 · 4:3 · 1:1 · 3:4 · 9:16 · 21:9 · adaptive | adaptive |
| Output Format | select | mp4 by default; mov encodes yuv444p for higher color fidelity in multi-round editing pipelines.mp4 · mov | mp4 |
| Generate Audio | select | Synchronized audio — voice, sound effects, and background music — generated with the video.false · true | true |
Pricing
From $0.14 per second, audio included
Pay per generation — only pay for what you use.
480p
$0.14/s
per second
720p
$0.30/s
per second
What a video costs
| Duration | 480p | 720p |
|---|---|---|
| 5 seconds | $0.70 | $1.51 |
| 10 seconds | $1.40 | $3.02 |
| 30 seconds | $4.20 | $9.03 |
Use cases
Film & storytelling
- Complete 30-second scenes
- Dialogue clips with lip-sync
- Establishing shots
- Previz and animatics
Marketing & ads
- Product spots with music
- UGC-style ads
- Brand stings
- Concept testing
Social content
- Vertical 9:16 stories
- 30-second narratives
- Sound-on hooks
- Series with consistent style
E-commerce
- Product showcase clips
- Lifestyle b-roll
- Seasonal campaign variants
- Listing videos
Frequently asked questions
What is Seedance 2.5 Reference to Video?+
It's the multimodal mode of ByteDance's Seedance 2.5: instead of starting from just a prompt, you guide the generation with your own reference material — images for characters and style, videos for motion, audio for pacing and lip-sync. Cite each asset in the prompt as @Image1, @Video1, @Audio1. On Scenetra it runs as a node in the visual workflow editor.
How many reference files can I use?+
Up to 30 reference images, up to 10 reference videos with a combined duration of 30 seconds, and up to 10 audio tracks, also 30 seconds combined — up to 50 assets per generation. Each image can be up to 30MB (300-6000px per side); videos are mp4/mov clips of 2-30 seconds each; audio is wav or mp3.
What is audio-only referencing?+
A Seedance 2.5 exclusive: you can drive a generation with sound alone. Supply one music track, voice recording, or sound effect — no reference images required — and the model generates visuals that follow it: pacing matches the beat, and speech produces lip-synced performance. It's the fastest way to turn a song or voiceover into video.
How do I reference my uploads in the prompt?+
Use @-syntax in submission order: @Image1 is your first image, @Video1 your first video, @Audio1 your first audio track. For example: 'The character from @Image1 walks through the city in @Image2, moving like @Video1, scored by @Audio1.' The model maps each mention to the exact asset.
Can Seedance 2.5 extend or edit an existing video?+
You can supply existing footage as a reference video and prompt a continuation or a stylistic match — that covers most extension-style use cases. The model's dedicated frame-level video-editing mode is not yet exposed on Scenetra.
How do I keep a character consistent across multiple videos?+
Reuse the same reference images across generations. Upload your character from a few angles, cite them as @Image1-@Image3 in every prompt, and the model keeps the identity stable from clip to clip — which is how episodic series and brand characters stay recognizable.
Can I use photos of real people as references?+
Directly uploading images or videos containing real human faces is not permitted by the model. Use model-generated characters, preset digital characters, or properly authorized real-person assets instead.
Related models
Seedance 2
ByteDance's Seedance 2.0 generates cinematic videos from text prompts with synchronized audio — voice, sound effects, and background music. Pick any duration from 4 to 15 seconds and resolutions up to 1080p, with per-second pricing that keeps 480p drafting cheap.
View model →Seedance 2 Mini
ByteDance's Seedance 2 Mini is the fast, lower-cost tier of Seedance 2 — generating cinematic videos from text prompts with synchronized audio. Mini delivers the same audio-visual sync at a fraction of the price, ideal for rapid iteration and high-volume social content.
View model →Kling 3.0 Pro
Kuaishou's flagship video generation model delivering stunning visual fidelity. Kling 3.0 Pro supports text-to-video, image-to-video with start/end frames, element references for character consistency, and native audio generation.
View model →Start creating with Seedance 2.5 Reference to Video
Use Seedance 2.5 Reference to Video alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started