Text & Image to VideoReference to Video

Consistent Characters from Reference Images — Veo 3.1

Google Veo 3.1 Reference-to-Video — the mode Google calls Ingredients to Video — generates videos with a consistent subject from up to 3 reference images. Keep a character, product, or style stable across shots, with native audio and 720p or 1080p output.

What makes Veo 3.1 Reference-to-Video different

One subject, consistent across every shot

Upload up to 3 reference images of a character, product, or animal — Google calls them 'ingredients' — and Veo 3.1 keeps that subject's appearance stable in the generated video. Reuse the same references across generations and the identity carries from clip to clip, which is the foundation for episodic content and recurring brand characters.

References guide identity, the prompt directs the scene

Unlike image-to-video, your reference images are not the first frame. They tell the model what the subject looks like, while your text prompt freely stages the scene — new setting, new camera move, new action — without being anchored to the composition of any input image.

Native audio in reference mode too

Reference-to-video keeps Veo 3.1's audio synthesis: dialogue, ambience, and effects generated in sync with the picture. Audio is on by default; turn it off and the per-second price drops from $0.40 to $0.20.

Generate your references on the same board

On Scenetra, character sheets are one node away: generate stills with an image model like Nano Banana 2, wire them straight into the Reference Images input, and iterate on the character and the video in one canvas — no exporting and re-uploading between tools.

Playground

Veo 3.1 Reference-to-Video
playground

A timelapse of a flower blooming in a sunlit meadow, cinematic quality

Drop images or click to upload

16:9
8s
1080p
true
Generate
Resultpreview
Sample output from Veo 3.1 Reference-to-Video

Try Veo 3.1 Reference-to-Video in Scenetra

Open Playground

Parameters

ParameterTypeDescriptionDefault
Prompt*textText prompt describing the video to generate around your reference subject.
Reference Images*image upload1 to 3 reference images for subject-consistent generation. Works with 16:9 aspect ratio and 8-second duration.
Aspect RatioselectVideo aspect ratio — reference-to-video works with 16:9.16:9 · 9:1616:9
DurationselectVideo duration — reference-to-video generates 8-second clips.8s
ResolutionselectResolution of the generated video.720p · 1080p1080p
Generate AudioselectGenerate native synchronized audio with the video — cheaper when disabled.false · truetrue
Negative PrompttextOptional description of what to exclude from the generated video.

Pricing

From $0.20 per second

Pay per generation — only pay for what you use.

Without audio

$0.20/s

per second

With audio

$0.40/s

per second

Use cases

Characters & series

  • Recurring characters across episodes
  • Brand mascots
  • AI influencer clips
  • Character-driven shorts

Product & e-commerce

  • Same product in new scenes
  • Lifestyle shots from packshots
  • Campaign variants
  • Product demos with audio

Film & previz

  • Cast consistency in previz
  • Costume and prop continuity
  • Scene exploration around one subject
  • Pitch footage from concept art

Social content

  • A persona that looks the same weekly
  • Pet and creator content
  • Branded skits
  • Sound-on character hooks

Related models

Frequently asked questions

What is Veo 3.1 Reference-to-Video (Ingredients to Video)?+

It's the reference-driven mode of Google's Veo 3.1 — Google markets it as 'Ingredients to Video'. Instead of starting from a first-frame image, you supply up to 3 reference images of a subject, and the model generates a new video in which that subject stays visually consistent while your prompt directs the scene. On Scenetra it runs as its own node in the visual workflow editor.

How many reference images can I use with Veo 3.1?+

1 to 3 reference images per generation. Fewer, cleaner references of the same subject tend to work better than mixing subjects — use multiple angles of one character or product rather than three different things.

What's the difference between reference images and image-to-video?+

Image-to-video uses your image as the literal first frame — the video starts from that exact composition. Reference-to-video uses your images as identity guides: the subject looks the same, but the framing, setting, and action come from your prompt. Use image-to-video to animate a specific shot; use reference-to-video to put a consistent subject into new shots.

How much does Veo 3.1 Reference-to-Video cost?+

The same per-second rate as standard Veo 3.1 at HD resolutions: $0.20 per second without audio or $0.40 per second with audio. Clips are 8 seconds long. You pay per generation from your Scenetra credit balance — no subscription required.

What are the limits of reference mode?+

Reference-to-video generates 8-second clips and works with the 16:9 aspect ratio, at 720p or 1080p (no 4K in this mode). If you need other durations, vertical video, or 4K, use the standard Veo 3.1 node's text- and image-to-video modes instead.

Can I try Veo 3.1 Reference-to-Video for free?+

You can try it without paying upfront: new Scenetra accounts get free welcome credits and a 7-day trial, which covers your first generations. After that it's pay-per-generation — there's no plan to subscribe to and nothing to cancel, and Scenetra runs in the browser on any machine.

Start creating with Veo 3.1 Reference-to-Video

Use Veo 3.1 Reference-to-Video alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.

Get Started Free