Google Veo 3.1 — Next-Gen AI Video Generation
Google's premier video generation model supporting text-to-video, image-to-video, and first-last-frame video creation. Veo 3.1 delivers up to 4K resolution with native audio synthesis, making it one of the most versatile video models available.
What makes Veo 3.1 different
Native audio, generated with the picture
Veo 3.1 synthesizes dialogue, ambient sound, and effects together with the video — no separate audio pass, no lip-sync post-processing. Audio is a toggle, and it's honest about cost: at 720p or 1080p you pay $0.40 per second with audio and $0.20 without, so silent drafts cost half as much as finished sound-on shots.
Up to 4K resolution
Most AI video models stop at 1080p. Veo 3.1 renders at 720p, 1080p, or full 4K, which holds up on large screens, in client presentations, and as source footage for further editing. 4K is billed at a higher per-second rate, so you can iterate cheaply at 720p and re-render the keeper at 4K.
First-last-frame control
Beyond plain text-to-video and image-to-video, Veo 3.1 accepts both a first and a last frame and generates the motion that connects them. That turns transformations, match cuts, and camera moves that must land on an exact composition into something you direct rather than reroll for.
A whole family on one board
Veo 3.1 comes in three price tiers on Scenetra — standard, Fast, and Lite — plus a Reference-to-Video mode for subject-consistent characters. They all run as nodes on the same canvas, next to Sora, Kling, and Seedance, so you can send one prompt through several models and keep the best take.
Playground
A timelapse of a flower blooming in a sunlit meadow, cinematic quality
Drop images or click to upload
Try Veo 3.1 in Scenetra
Open PlaygroundParameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Text describing the video you want to generate. | — |
| Duration | select | Length of the generated video.4s · 6s · 8s | 8s |
| Resolution | select | Output video resolution — higher resolutions cost more.720p · 1080p · 4K | 720p |
| Aspect Ratio | select | Video aspect ratio.16:9 · 9:16 | 16:9 |
| First Frame | image upload | Optional starting frame for image-to-video generation. | — |
| Last Frame | image upload | Optional ending frame for first-last-frame video mode. | — |
| Generate Audio | select | Generate native synchronized audio with the video — cheaper when disabled.false · true | true |
Pricing
From $0.20 per second
Pay per generation — only pay for what you use.
720p/1080p
$0.20/s
per second
With audio
$0.40/s
per second
4K
$0.40/s
per second
4K + audio
$0.60/s
per second
What a video costs
| Duration | 720p, audio on | 1080p, audio on |
|---|---|---|
| 5 seconds | $2.00 | $2.00 |
| 10 seconds | $4.00 | $4.00 |
Use cases
Film & storytelling
- Cinematic establishing shots
- Dialogue scenes with native audio
- First-to-last-frame match cuts
- Previz and pitch reels
Marketing & ads
- Product spots in 4K
- Sound-on brand films
- Concept testing before a shoot
- Localized ad variants
Social content
- Vertical 9:16 clips
- Sound-on hooks
- Animate a hero image
- Weekly series from one style
Agencies & studios
- Client boards with multiple takes
- 4K masters for delivery
- Draft on Fast, finish on standard
- Side-by-side model comparisons
Related models
Veo 3.1 Fast
Google Veo 3.1 Fast generates high-quality videos from text or images at a fraction of the cost of standard Veo 3.1. Supports text-to-video, image-to-video, and first-last-frame mode with native audio at up to 4K resolution.
View model →Veo 3.1 Lite
Google Veo 3.1 Lite is the most affordable Veo tier, optimized for high-volume video generation. Generates 720p or 1080p videos from text or images with optional synchronized audio.
View model →Seedance 2.5
ByteDance's Seedance 2.5 generates cinematic videos up to 30 seconds from text prompts with synchronized audio — voice, sound effects, and background music — built in by default. The newest Seedance generation, with mp4 or high-fidelity mov output for multi-round editing pipelines.
View model →Kling 3.0 Pro
Kuaishou's flagship video generation model delivering stunning visual fidelity. Kling 3.0 Pro supports text-to-video, image-to-video with start/end frames, element references for character consistency, and native audio generation.
View model →Frequently asked questions
What is Veo 3.1 and how can I use it?+
Veo 3.1 is Google's flagship AI video model. It generates videos from text prompts or input images, with native synchronized audio and resolutions up to 4K. On Scenetra you use it without an API key or waitlist: sign up, drop a Veo 3.1 node onto your board, write a prompt (or connect an image), and generate. Scenetra is web-based, so it works anywhere.
How much does Google Veo 3.1 cost?+
On Scenetra, Veo 3.1 is billed per second of video: $0.20/s without audio or $0.40/s with audio at 720p or 1080p, and $0.40/s without audio or $0.60/s with audio at 4K. A 5-second 1080p clip with audio costs about $2.00. You only pay for what you generate.
Is Veo 3.1 completely free?+
No — Veo 3.1 is a paid model wherever you run it. On Scenetra you can still try it without paying: new accounts get free welcome credits and a 7-day trial that covers your first generations. After that it's pay-per-generation from your credit balance, with no subscription gating the model and nothing to cancel.
Do I need a monthly Google AI subscription to use Veo 3.1?+
Not on Scenetra. There is no monthly plan required to access Veo 3.1 — you pay per second of generated video from a credit balance. That's usually cheaper than a subscription if you generate in bursts, and there's no tier that locks Veo behind a higher plan.
Which is better, Veo 3.1 or Sora 2?+
They trade blows. Veo 3.1's edge is native audio at every tier, resolutions up to 4K, and first-last-frame control; Sora 2 and Sora 2 Pro are strong on prompt adherence and physicality, and many creators prefer them for certain styles. On Scenetra both run side by side on the same board, so the practical answer is to send your prompt through each and compare.
Does Veo 3.1 generate audio?+
Yes — Veo 3.1 generates dialogue, ambient sound, and effects natively, synchronized with the picture. Audio is on by default and can be toggled off, which halves the per-second price at 720p/1080p ($0.20/s instead of $0.40/s).
Can Veo 3.1 use reference images (ingredients)?+
Yes — what Google calls 'Ingredients to Video' is available on Scenetra as Veo 3.1 Reference-to-Video, a separate mode of this model. You upload up to 3 reference images and the model keeps that subject consistent in the generated video. The standard node covers text-to-video, image-to-video (first frame), and first-last-frame instead.
What durations, resolutions, and aspect ratios does Veo 3.1 support?+
Clips of 4, 6, or 8 seconds; resolutions of 720p, 1080p, or 4K; and 16:9 or 9:16 aspect ratios. You can start from text only, a first-frame image, or a first and last frame pair.
Start creating with Veo 3.1
Use Veo 3.1 alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started Free