Reference-to-Video with PixVerse v6
Drive PixVerse v6 with up to 7 reference images and bind them directly in your prompt as @image1, @image2. Keep characters, objects and scenes consistent across 1-15 second videos at 360p to 1080p, with eight aspect ratios and optional synchronized audio. From $0.0225 per second.
What makes PixVerse v6 different
Up to 7 references, bound by name
Upload up to 7 images of characters, objects or scenes, then address them directly in the prompt by upload order: "@image1 walks through @image2 holding @image3". Instead of hoping the model picks up the right subject, you tell it exactly which reference plays which role in the shot.
Consistent characters across a whole series
Because the references travel with every generation, the same character can appear across dozens of clips — new scenes, new actions, same face and outfit. Combine references freely: a character from one image, a prop from another, a location from a third, all composed into one 1-15 second video.
Full v6 quality and format range
Reference-to-Video keeps everything else v6 offers: four resolutions from 360p to 1080p, all eight aspect ratios including 9:16 and 21:9, 1-15 second durations, and optional synchronized audio with BGM, SFX and dialogue as a per-second surcharge.
A small premium over text-to-video
Reference mode runs 12.5% above the standard v6 rates — from $0.0225/s at 360p to $0.081/s at 1080p, with audio adding $0.009-$0.0225/s. A 5-second 720p reference video is about $0.20. On Scenetra it's billed at provider cost with 0% markup, and free welcome credits cover your first tests.
Playground
A timelapse of a flower blooming in a sunlit meadow, cinematic quality
Drop images or click to upload
Try PixVerse v6 in Scenetra
Open PlaygroundParameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Describe the video and bind your references by upload order: "@image1 walks through @image2". With audio enabled, sound is generated to match. | — |
| Reference Images | image upload | 1-7 reference images for characters, objects or scenes. Bind them in the prompt as @image1, @image2, ... (in upload order). | — |
| Duration | select | Duration of the generated video in seconds (1-15).1 · 2 · 3 · 4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 | 5 |
| Resolution | select | Output resolution. 720p is the model's default.360p · 540p · 720p · 1080p | 720p |
| Aspect Ratio | select | Aspect ratio of the generated video.16:9 · 4:3 · 1:1 · 3:4 · 9:16 · 2:3 · 3:2 · 21:9 | 16:9 |
| Generate Audio | boolean | Generate synchronized audio (BGM, SFX, dialogue). Adds a per-second surcharge. | false |
Pricing
From $0.0225 per second
Pay per generation — only pay for what you use.
360p
$0.0225/s
per second
540p
$0.0315/s
per second
720p
$0.0405/s
per second
1080p
$0.081/s
per second
Synchronized audio
+$0.009-$0.0225/s
per second
What a video costs
| Duration | 360p | 540p | 720p | 1080p |
|---|---|---|---|---|
| 5 seconds | $0.11 | $0.16 | $0.20 | $0.41 |
| 10 seconds | $0.23 | $0.32 | $0.41 | $0.81 |
Use cases
Character content
- Recurring characters across episodes
- Brand mascot videos
- Virtual influencer clips
- Cast multiple characters in one scene
Product & e-commerce
- Same product in many settings
- Model wearing your exact item
- Prop-accurate demo videos
- Catalog-consistent campaigns
Film & storyboarding
- Cast consistency across shots
- Location continuity from scene refs
- Costume-accurate previz
- Character-driven mood pieces
Social series
- Episodic character formats
- Consistent avatar skits
- Fan-art character animations
- Serialized story clips
Related models
Seedance 2.5 Reference-to-Video
Generate videos from up to 30 reference images, 10 reference videos, and 10 audio tracks with Seedance 2.5 by ByteDance. Cite references directly in your prompt as @Image1, @Video1, @Audio1. Unique to 2.5: audio-only referencing — a single music or voice track can drive visual pacing, beat matching, and lip-sync.
View model →Wan 3.0
Generate video from up to 10 reference images with Alibaba's Wan 3.0. The model carries the people, objects, and styles from your references into a continuous take of 2 to 30 seconds, with native synchronized audio generated in the same pass. Output at 480p, 720p, or 1080p — reference images are free, and audio adds nothing to the price.
View model →MiniMax H3 Reference-to-Video
Direct MiniMax H3 (Hailuo 03) with the material you already have: up to 9 reference images, 3 reference video clips, and 3 audio tracks in a single generation. Keep subjects and products consistent across shots, borrow motion and pacing from reference footage, sync action to a music track, and get native audio in the output — 5-15 seconds, from 480P up to 4K.
View model →Grok Imagine Video 1.5
Generate videos from up to 7 reference images with xAI's Grok Imagine Video 1.5. Feed it the people, objects, and styles you want on screen, cite them in the prompt as @Image1, @Image2, and get a clip of up to 15 seconds with synchronized audio — consistent characters, no fine-tuning.
View model →Frequently asked questions
What is PixVerse v6 Reference-to-Video?+
A v6 mode that generates video from up to 7 reference images. You upload references for characters, objects or scenes and bind them in the prompt by upload order — "@image1 rides @image2 through @image3" — and the model keeps those subjects consistent in the output. Videos run 1-15 seconds at 360p to 1080p in eight aspect ratios.
How do I reference images in the prompt?+
By upload order: the first image you add is @image1, the second is @image2, and so on up to @image7. Write the prompt as if the tags were the subjects — "@image1 walks through @image2" — so each reference gets an explicit role instead of being blended in loosely.
How much does PixVerse v6 reference-to-video cost?+
From $0.0225 per second at 360p, $0.0315/s at 540p, $0.0405/s at 720p and $0.081/s at 1080p — a 12.5% premium over v6's text-to-video rates. A 5-second 720p clip is about $0.20 and a 10-second 1080p clip about $0.81. Synchronized audio adds $0.009-$0.0225 per second. Scenetra bills at provider cost with 0% markup.
How is this different from image-to-video?+
Image-to-video uses one image as the literal first frame and animates from it. Reference-to-video uses your images as identity references — they define what characters, objects and scenes look like, while the prompt freely composes a new shot around them in any of the eight aspect ratios. Use i2v to animate a finished frame; use references to keep subjects consistent across many different shots.
Can I keep the same character across multiple videos?+
Yes — that's the core use case. Reuse the same reference image (or set of images) across generations and the character stays consistent scene after scene. On a Scenetra board you can fan one reference image out to many Reference to Video nodes and batch a whole series.
Does reference-to-video support audio?+
Yes. Enable the Generate Audio toggle and v6 generates BGM, sound effects and dialogue synchronized to the video, billed as a per-second surcharge on top of the base rate.
Start creating with PixVerse v6
Use PixVerse v6 alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started Free