Text-to-Video with Native Audio using Grok Imagine Video 1.5
xAI's Grok Imagine Video 1.5 turns a text prompt into a video with synchronized audio generated in the same pass — no separate sound step and no audio surcharge. Choose any duration from 1 to 15 seconds, three resolutions from 480p to 1080p, and eight aspect ratios from widescreen 16:9 to vertical 9:16.
What makes Grok Imagine Video 1.5 different
Native audio, included in the price
Grok Imagine Video 1.5 generates synchronized audio with every clip in the same pass as the picture, and it's covered by the per-second rate rather than billed as an add-on. You can even direct the soundtrack from the prompt — name the ambience, effects, or mood you want and the model scores the scene to match what happens on screen.
Free-form durations from 1 to 15 seconds
Pick any length from 1 to 15 seconds instead of choosing between a couple of fixed presets. One- and two-second clips make cheap motion tests, six seconds is the sweet-spot default, and 15 seconds is enough for a full beat of action with sound — at 480p, a 15-second draft still costs well under $0.20.
Three resolutions, one cheap draft rung
480p runs at $0.012 per second, 720p at $0.0225, and 1080p at $0.04 — so a 5-second draft costs $0.06 and a 5-second full-HD final $0.20. A practical Scenetra workflow: iterate your prompt at 480p until motion and audio land, then flip the resolution select to 1080p for the final render. Same node, same wiring, and every generation is billed at provider cost with 0% markup.
Eight aspect ratios, or let the model decide
Grok Imagine Video 1.5 outputs 16:9, 4:3, 3:2, 1:1, 2:3, 3:4, and 9:16 — covering widescreen, square, and vertical social formats — or you can leave the ratio on Auto and let the model pick a frame that suits your prompt. That makes it easy to render the same idea as both a landscape hero clip and a 9:16 story cut.
Playground
A timelapse of a flower blooming in a sunlit meadow, cinematic quality
Try Grok Imagine Video 1.5 in Scenetra
Open PlaygroundParameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Describe the video: subject, motion, camera work, and audio direction. Native synchronized audio is generated in the same pass. | — |
| Resolution | select | Output resolution. 480p is the budget rung; 1080p is the premium rung.480p · 720p · 1080p | 720p |
| Aspect Ratio | select | Aspect ratio of the generated video. Auto lets the model choose from your prompt.Auto · 16:9 · 4:3 · 3:2 · 1:1 · 2:3 · 3:4 · 9:16 | auto |
| Duration | select | Length of the generated video in seconds (1-15).1 · 2 · 3 · 4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 | 6 |
Pricing
From $0.012 per second
Pay per generation — only pay for what you use.
480p
$0.012/s
per second
720p
$0.0225/s
per second
1080p
$0.04/s
per second
What a video costs
| Duration | 480p | 720p | 1080p |
|---|---|---|---|
| 5 seconds | $0.06 | $0.113 | $0.20 |
| 10 seconds | $0.12 | $0.225 | $0.40 |
Use cases
Social content
- Sound-on hooks and openers
- Vertical 9:16 story clips
- Meme and reaction videos
- Daily posting at draft prices
Marketing & ads
- Concept ads with audio
- Product teaser clips
- A/B test multiple ad variants
- Seasonal campaign spots
Film & previz
- Establishing shots from a prompt
- Mood pieces with ambience
- Cheap 480p shot exploration
- Scene beats up to 15 seconds
Music & audio-led
- Visualizers scored by the model
- Ambient loops with sound
- Lyric-inspired scene clips
- Sound-design experiments
Related models
Seedance 2.5
ByteDance's Seedance 2.5 generates cinematic videos up to 30 seconds from text prompts with synchronized audio — voice, sound effects, and background music — built in by default. The newest Seedance generation, with mp4 or high-fidelity mov output for multi-round editing pipelines.
View model →Kling 3.0 Pro
Kuaishou's flagship video generation model delivering stunning visual fidelity. Kling 3.0 Pro supports text-to-video, image-to-video with start/end frames, element references for character consistency, and native audio generation.
View model →PixVerse v6
PixVerse's flagship video model generates 1-15 second clips from a text prompt with optional synchronized audio, multi-shot sequences, four resolutions from 360p to 1080p, and eight aspect ratios from vertical 9:16 to cinematic 21:9. Starting at $0.02 per second, it is one of the most affordable ways to produce finished AI video.
View model →Happy Horse 1.1
Alibaba's Happy Horse 1.1 generates 3-15 second videos with native audio straight from a text prompt. Dialogue, music, and sound effects are directed through the same prompt as the visuals, and nine aspect ratios from 9:21 to 21:9 cover everything from vertical stories to ultrawide cinematics at 480p, 720p, or 1080p.
View model →Frequently asked questions
What is Grok Imagine Video 1.5?+
It's xAI's latest video generation model. Give it a text prompt describing the subject, motion, camera work, and even the audio direction, and it generates a video of 1 to 15 seconds with synchronized sound created in the same pass. On Scenetra it runs as a node in the visual workflow editor, alongside image-to-video and reference-to-video modes of the same model.
How much does Grok Imagine Video 1.5 cost?+
$0.012 per second at 480p, $0.0225 per second at 720p, and $0.04 per second at 1080p. That's $0.06 for a 5-second 480p draft and $0.40 for a 10-second 1080p final. Audio is included in those rates, and Scenetra bills every generation at provider cost with 0% markup.
Does Grok Imagine Video 1.5 generate audio?+
Yes — synchronized audio is generated together with the video in a single pass, and it's included in the per-second price rather than billed as a separate add-on. You can steer it from the prompt: describe the ambience, sound effects, or mood you want alongside the visual direction.
Can I try Grok Imagine free online?+
Yes — new Scenetra accounts get free welcome credits and a 7-day trial, which covers plenty of first generations given a 5-second 480p clip costs about $0.06. After that it's pay-per-generation with no subscription gate and nothing to cancel.
What resolutions and aspect ratios does it support?+
Three resolutions — 480p, 720p, and 1080p — and eight aspect ratio options: 16:9, 4:3, 3:2, 1:1, 2:3, 3:4, 9:16, or Auto, which lets the model choose a frame based on your prompt. 1080p is available in text-to-video and costs $0.04 per second.
What's the difference between the text-to-video, image-to-video, and reference-to-video modes?+
Text-to-video (this page) generates a clip from a prompt alone. Image-to-video animates one uploaded still into motion. Reference-to-video takes up to 7 reference images and keeps their people, objects, or styles consistent in the generated video. All three share the same 1-15 second durations and native audio, and all run as nodes on the same Scenetra board.
How should I prompt Grok Imagine Video 1.5?+
Be concrete about four things: the subject, the motion ('she turns toward the camera'), the camera work ('slow dolly in'), and the audio ('distant thunder, soft rain'). Short, specific prompts tend to beat long scene descriptions, and at $0.012 per second for 480p it's cheap to iterate on a Scenetra board until the motion lands.
Start creating with Grok Imagine Video 1.5
Use Grok Imagine Video 1.5 alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started Free