Reference-to-Video with Grok Imagine Video 1.5
Generate videos from up to 7 reference images with xAI's Grok Imagine Video 1.5. Feed it the people, objects, and styles you want on screen, cite them in the prompt as @Image1, @Image2, and get a clip of up to 15 seconds with synchronized audio — consistent characters, no fine-tuning.
What makes Grok Imagine Video 1.5 different
Up to 7 references, cited by name
Upload between 1 and 7 reference images — a character, a product, a location, a style frame — and address them directly in the prompt as @Image1, @Image2, and so on. That gives you precise control over who does what: '@Image1 picks up @Image2 and walks toward the window' reads exactly the way the model executes it.
Consistent people, objects, and styles
Reference-to-video is how you keep the same face, the same product, or the same visual style across many generated clips without training anything. Generate a character sheet with an image model on the same Scenetra board, wire those stills into this node, and every shot in your sequence features the same character.
Native audio, included in the price
Like the rest of the Grok Imagine Video 1.5 family, reference-to-video generates synchronized audio in the same pass as the picture, covered by the per-second rate. Describe the ambience or effects you want in the prompt and the soundtrack matches the action on screen.
Draft at 480p, ship at 720p
Reference mode offers 480p at $0.012 per second and 720p at $0.0225 — a 5-second draft is $0.06 and a 10-second 720p final about $0.23; 1080p is not available with references. Reference images may add a small flat fee of about $0.01 each. Iterate cheap at 480p until the character reads right, then switch the select to 720p — billed at provider cost with 0% markup.
Playground
A timelapse of a flower blooming in a sunlit meadow, cinematic quality
Drop images or click to upload
Try Grok Imagine Video 1.5 in Scenetra
Open PlaygroundParameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Describe the video: subject, motion, camera work, and audio direction. Native synchronized audio is generated in the same pass. | — |
| Reference Images* | image upload | 1-7 reference images contributing people, objects, or styles. Cite them in the prompt as @Image1, @Image2, and so on. | — |
| Resolution | select | Output resolution. 480p is the budget rung; 720p is the production rung.480p · 720p | 480p |
| Aspect Ratio | select | Aspect ratio of the generated video. Auto lets the model choose from your prompt.Auto · 16:9 · 4:3 · 3:2 · 1:1 · 2:3 · 3:4 · 9:16 | auto |
| Duration | select | Length of the generated video in seconds (1-15).1 · 2 · 3 · 4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 | 6 |
Pricing
From $0.012 per second
Pay per generation — only pay for what you use.
480p
$0.012/s
per second
720p
$0.0225/s
per second
What a video costs
| Duration | 480p | 720p |
|---|---|---|
| 5 seconds | $0.06 | $0.113 |
| 10 seconds | $0.12 | $0.225 |
Use cases
Characters & storytelling
- Same character across every shot
- Multi-scene story sequences
- Character and prop interactions
- Style-matched episode clips
Product & e-commerce
- Your exact product in motion
- Product plus model scenes
- Brand-style campaign clips
- Variant videos from one shoot
Social & creators
- Recurring mascot content
- Avatar clips with sound
- Fan art brought to life
- Series with a consistent look
Film & previz
- Cast continuity in previz
- Wardrobe and prop consistency
- Location-matched shots
- Style-frame-driven sequences
Related models
Seedance 2.5 Reference-to-Video
Generate videos from up to 30 reference images, 10 reference videos, and 10 audio tracks with Seedance 2.5 by ByteDance. Cite references directly in your prompt as @Image1, @Video1, @Audio1. Unique to 2.5: audio-only referencing — a single music or voice track can drive visual pacing, beat matching, and lip-sync.
View model →Wan 3.0
Generate video from up to 10 reference images with Alibaba's Wan 3.0. The model carries the people, objects, and styles from your references into a continuous take of 2 to 30 seconds, with native synchronized audio generated in the same pass. Output at 480p, 720p, or 1080p — reference images are free, and audio adds nothing to the price.
View model →PixVerse v6
Drive PixVerse v6 with up to 7 reference images and bind them directly in your prompt as @image1, @image2. Keep characters, objects and scenes consistent across 1-15 second videos at 360p to 1080p, with eight aspect ratios and optional synchronized audio. From $0.0225 per second.
View model →Frequently asked questions
What is reference-to-video in Grok Imagine Video 1.5?+
It's the mode of xAI's Grok Imagine Video 1.5 that generates a video from up to 7 reference images instead of a single starting frame. The references contribute people, objects, or styles — the model keeps them consistent in the output — and you cite them in the prompt as @Image1, @Image2 to direct who does what. Clips run 1 to 15 seconds with synchronized audio included.
How is reference-to-video different from image-to-video?+
Image-to-video animates one still image — the output starts from and looks like your upload. Reference-to-video composes a new scene featuring the people, objects, or styles from up to 7 images, so the framing and setting come from your prompt while the subjects stay consistent. Use image-to-video to animate a finished frame, and references to keep a character or product recurring across many different shots.
How much does Grok reference-to-video cost?+
$0.012 per second at 480p and $0.0225 per second at 720p — about $0.06 for a 5-second 480p draft and $0.23 for a 10-second 720p clip, with audio included in the rate. Reference images may add a small flat fee of about $0.01 each. 1080p is not available in reference mode.
How many reference images can I use?+
Between 1 and 7 per generation. Each can contribute something different — one for the main character, one for an object they hold, one for the visual style — and you reference them by position in the prompt as @Image1 through @Image7.
How do I keep a character consistent across multiple videos?+
Reuse the same reference images for every generation. On a Scenetra board you can generate a character with an image model like Nano Banana 2, wire that output into the Grok reference-to-video node, and render shot after shot with the same face — changing only the prompt. No fine-tuning or training step is involved.
Does reference-to-video generate audio?+
Yes — like the whole Grok Imagine Video 1.5 family, it generates synchronized audio with the video in a single pass, and the sound is included in the per-second price. Describe the ambience or effects you want in the prompt alongside the action.
Can I try it free online?+
Yes — Scenetra gives new accounts free welcome credits and a 7-day trial, and with 5-second 480p drafts at roughly $0.06 that covers plenty of experiments. After the trial it's pay-per-generation with no subscription required.
Start creating with Grok Imagine Video 1.5
Use Grok Imagine Video 1.5 alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started Free