Consistent AI Video with xAI Grok Imagine
Generate videos with consistent subject appearance from up to 7 reference images using xAI's Grok Imagine Video. Reference each image in your prompt as @Image1, @Image2, etc. for style and content guidance.
What makes Grok Reference-to-Video different
Up to 7 reference images per generation
Instead of animating a single first frame, Grok Reference-to-Video takes up to 7 images and uses them as guidance for who and what appears in the video. Upload a character from several angles, a product, or a style board, and the generation draws from all of them rather than being locked to one composition.
Cite references directly in the prompt
Each upload gets an @-handle in submission order: @Image1, @Image2, and so on. Write prompts like 'the character in @Image1 walks through the market from @Image3' and the model knows exactly which asset plays which role — no guessing from a pile of attachments.
Same budget pricing as Grok Video
Reference mode costs the same per-second rate as the rest of the family: $0.05 per second at 480p and $0.07 at 720p, with no per-image surcharge for references. A 10-second 720p generation with a full set of 7 references is still about $0.70 — cheap enough to iterate until the identity holds.
Built for recurring characters
Because the same reference set can be reused across generations, subjects keep their identity from clip to clip. Keep your character's reference images wired into the node on your Scenetra board and every new prompt produces on-model footage — the foundation for episodic series and brand characters.
Playground
A timelapse of a flower blooming in a sunlit meadow, cinematic quality
Drop images or click to upload
Try Grok Reference-to-Video in Scenetra
Open PlaygroundParameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Reference images in the prompt as @Image1, @Image2, etc. | — |
| Reference Images* | image upload | Up to 7 reference images for style and content guidance. | — |
| Duration | number | Video duration in seconds. | — |
| Aspect Ratio | select | Aspect ratio of the generated video.16:9 · 4:3 · 3:2 · 1:1 · 2:3 · 3:4 · 9:16 | — |
| Resolution | select | Resolution of the output video.480p · 720p | 480p |
Pricing
From $0.05 per second
Pay per generation — only pay for what you use.
480p
$0.05/s
per second
720p
$0.07/s
per second
What a video costs
| Duration | 480p | 720p |
|---|---|---|
| 5 seconds | $0.25 | $0.35 |
| 10 seconds | $0.50 | $0.70 |
Use cases
Characters & series
- Consistent recurring characters
- Episodic story clips
- Brand mascots
- AI influencer content
Brand & product
- On-brand product videos
- Multi-angle product motion
- Campaign variations
- Style-board-driven ads
Style transfer
- Match a reference look
- Moodboard-to-motion
- Scene-consistent shots
- Art style continuity
Social content
- Character-led shorts
- Vertical 9:16 series
- Consistent avatar clips
- Low-cost iteration runs
Related models
Seedance 2.5 Reference-to-Video
Generate videos from up to 30 reference images, 10 reference videos, and 10 audio tracks with Seedance 2.5 by ByteDance. Cite references directly in your prompt as @Image1, @Video1, @Audio1. Unique to 2.5: audio-only referencing — a single music or voice track can drive visual pacing, beat matching, and lip-sync.
View model →Seedance 2 Reference-to-Video
Generate videos from up to 9 reference images, 3 reference videos, and 3 audio tracks with Seedance 2.0 by ByteDance. Mention references in your prompt — 'image 1', 'video 1' — for character consistency, motion transfer, and audio-driven lip-sync, at resolutions up to 1080p.
View model →Happy Horse Reference-to-Video
Alibaba's Happy Horse reference-to-video generates videos from 1-9 reference images. Reference each subject in your prompt as character1, character2, ... character9 (order matches your uploaded images) for consistent multi-character scenes. 720p or 1080p output, 3-15 second durations.
View model →Grok Imagine Video v1.5
xAI's latest image-to-video model. Grok Imagine Video v1.5 animates a single input image into a video up to 15 seconds long with synchronized audio generated at no extra cost. Pick 480p for fast iteration or 720p for production-quality output.
View model →Frequently asked questions
What is Grok Reference to Video?+
It's the multi-reference mode of xAI's Grok Imagine Video: instead of starting from a text prompt or a single frame, you supply up to 7 reference images and cite them in your prompt as @Image1, @Image2, and so on. The model uses them for subject identity, style, and content guidance. On Scenetra it runs as a node in the visual workflow editor.
How many reference images can I use?+
Up to 7 per generation. Each is addressable in the prompt by its submission order — @Image1 is your first upload, @Image7 your last — so you can assign roles explicitly, like a character from @Image1 wearing the outfit in @Image2.
How is Reference to Video different from image to video?+
Image to video animates one image as the literal first frame of the clip. Reference to video treats your images as guidance rather than a starting frame: the model composes new shots that keep the subjects and style of your references. Use i2v when you want to control exactly what the opening frame looks like; use references when you want the same character or product across freely composed shots.
How much does Grok Reference to Video cost?+
The same as the other Grok Video mode: $0.05 per second at 480p and $0.07 per second at 720p, with no extra charge for reference images. A 5-second 480p test is about $0.25; a 10-second 720p clip about $0.70.
Can I try Grok Reference to Video for free?+
Yes — Scenetra runs in the browser and new accounts get free welcome credits plus a 7-day trial, which covers your first generations. After that it's pay-per-generation from your credit balance, with no subscription required.
How do I keep a character consistent across multiple videos?+
Reuse the same reference images for every generation. Upload your character from a few angles, cite them with the same @-handles in each prompt, and the identity stays stable from clip to clip. On a Scenetra board you can keep the references wired into the node and just change the prompt per shot.
Start creating with Grok Reference-to-Video
Use Grok Reference-to-Video alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started Free