Reference-to-Video with Gemini Omni Flash
Generate 3-10 second videos from up to 10 reference images with Google's Gemini Omni Flash. Keep subjects, products, and style consistent across shots while the model composes a new scene — with dialogue, music, and sound effects rendered natively and directed through the prompt.
What makes Gemini Omni Flash different
Up to 10 references, one consistent scene
Unlike image-to-video, which pins the first frame, reference-to-video treats your uploads as identity and style anchors: a character from several angles, a product, a wardrobe, a color script. The model composes a new scene that keeps those subjects consistent — the core tool for recurring characters and on-brand product footage.
Native audio, directed by prompt
Dialogue, music, and sound effects are generated in the same pass as the picture. Your referenced character doesn't just look consistent — they can speak a scripted line while the scene carries its own ambience and score, all written into a single prompt with no separate audio step.
Even durations hit the price floor
Pricing varies with duration, and longer can be cheaper: a 10-second 720p clip costs $0.50 total ($0.05 per second) while a 5-second one costs $0.65 — even durations (4, 6, 8, 10 seconds) get the floor rate at 720p. On 10-second clips, 1080p upscaling is $0.63 and 4K is $1.05.
Reference sets live on the board
On a Scenetra board, your reference images are nodes you wire in — generate character sheets with an image model, connect them to a Gemini Omni Flash reference-to-video node, and reuse the same set across many shots. Every generation is billed at provider cost with 0% markup, pay-per-generation.
Playground
A timelapse of a flower blooming in a sunlit meadow, cinematic quality
Drop images or click to upload
Try Gemini Omni Flash in Scenetra
Open PlaygroundParameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Describe the result and how the reference images should be used (subjects, style, composition). Audio is rendered natively and directed through this prompt. | — |
| Reference Images | image upload | Up to 10 reference images for subjects, products, style, or composition. | — |
| Duration | select | Duration of the video in seconds (3-10).3 · 4 · 5 · 6 · 7 · 8 · 9 · 10 | 8 |
| Resolution | select | 720p is native; 1080p and 4K upscale the 720p render.720p · 1080p · 4k | 720p |
| Aspect Ratio | select | Aspect ratio of the generated video.16:9 · 9:16 | 16:9 |
Pricing
From $0.05 per second
Pay per generation — only pay for what you use.
720p, 5s clip
$0.13/s
per second
720p, 10s clip
$0.05/s
per second
1080p, 10s clip
$0.063/s
per second
4K, 10s clip
$0.105/s
per second
What a video costs
| Duration | 720p | 1080p | 4K |
|---|---|---|---|
| 5 seconds | $0.65 | — | — |
| 10 seconds | $0.50 | $0.63 | $1.05 |
Use cases
Characters & storytelling
- Recurring characters across shots
- Multi-scene story sequences
- Character dialogue scenes
- Style-consistent episodes
Product & e-commerce
- Product videos from catalog photos
- On-brand campaign footage
- Product-in-scene compositions
- Consistent product families
Brand & advertising
- Mascots that stay on-model
- Brand-style locked footage
- Spokesperson-style clips
- Campaign variants from one set
Fashion & lifestyle
- Outfit references in motion
- Lookbook video clips
- Consistent model appearances
- Styled scene compositions
Related models
Wan 3.0
Generate video from up to 10 reference images with Alibaba's Wan 3.0. The model carries the people, objects, and styles from your references into a continuous take of 2 to 30 seconds, with native synchronized audio generated in the same pass. Output at 480p, 720p, or 1080p — reference images are free, and audio adds nothing to the price.
View model →MiniMax H3 Reference-to-Video
Direct MiniMax H3 (Hailuo 03) with the material you already have: up to 9 reference images, 3 reference video clips, and 3 audio tracks in a single generation. Keep subjects and products consistent across shots, borrow motion and pacing from reference footage, sync action to a music track, and get native audio in the output — 5-15 seconds, from 480P up to 4K.
View model →PixVerse v6
Drive PixVerse v6 with up to 7 reference images and bind them directly in your prompt as @image1, @image2. Keep characters, objects and scenes consistent across 1-15 second videos at 360p to 1080p, with eight aspect ratios and optional synchronized audio. From $0.0225 per second.
View model →Seedance 2.5 Reference-to-Video
Generate videos from up to 30 reference images, 10 reference videos, and 10 audio tracks with Seedance 2.5 by ByteDance. Cite references directly in your prompt as @Image1, @Video1, @Audio1. Unique to 2.5: audio-only referencing — a single music or voice track can drive visual pacing, beat matching, and lip-sync.
View model →Frequently asked questions
What is reference-to-video in Gemini Omni Flash?+
It's a generation mode that takes up to 10 reference images — subjects, products, styles, compositions — and composes a new 3-10 second video that keeps them consistent. Unlike image-to-video, the references aren't the first frame; they're identity anchors the model draws from while building a new scene, with native audio directed through the prompt.
How is reference-to-video different from image-to-video?+
Image-to-video starts the clip from your exact uploaded frame and animates forward. Reference-to-video composes a brand-new scene while keeping the referenced subjects and style consistent — better when you want the same character or product in a different setting, angle, or action than your source photos show.
How many reference images can I use?+
Up to 10 per generation. Use them for whatever needs to stay consistent: a character from several angles, a product, a wardrobe, a location, or a visual style. The prompt tells the model how each reference should be used in the final scene.
How much does Gemini Omni Flash reference-to-video cost?+
A 10-second 720p clip costs $0.50 ($0.05 per second) and a 5-second one $0.65 — pricing varies with duration, and even durations (4, 6, 8, 10 seconds) get the cheapest rate. On 10-second clips, 1080p costs $0.63 and 4K $1.05. All billed at provider cost with 0% markup.
Does it generate audio too?+
Yes — like every Gemini Omni Flash mode, dialogue, music, and sound effects are rendered natively in the same pass as the picture and directed through the prompt. Your referenced character can deliver a scripted line in the generated scene.
Can I try it free?+
Yes — new Scenetra accounts get free welcome credits and a 7-day trial that cover your first generations. After that it's pay-per-generation from your credit balance, with no subscription gate.
Start creating with Gemini Omni Flash
Use Gemini Omni Flash alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started Free