Grok Imagine Video 1.5 Image to Video
xAI's latest image-to-video model. Grok Imagine Video v1.5 animates a single input image into a video up to 15 seconds long with synchronized audio generated at no extra cost. Pick 480p for fast iteration, or 720p and 1080p for production-quality output.
84% less than the priciest provider (7 compared). Same model, same output.
Starting frame. The video animates from this image.
Resolution
Output resolution. 480p is the budget rung; 1080p is the premium rung.
Duration
Length of the generated video in seconds (1-15). Billed per second.
84% less than the priciest provider (7 compared). Same model, same output.
Your result renders here.
Use Grok Imagine Video 1.5 Image to Video via the API
Generate with Grok Imagine Video 1.5 Image to Video from your own code. Async by design: submit a generation, get a run id back instantly, poll until it completes. Billed in credits at the best available provider price — from $0.012 per second.
1 — Start a generation
curl -X POST https://app.scenetra.com/api/v1/generate/grok-imagine-video-1.5-i2v \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"prompt":"Slow push-in, the subject turns to camera and smiles, soft wind in the hair, gentle ambient sound","image_url":"https://example.com/start-frame.jpg","resolution":"720p","duration":"6"}'Response (202)
{
"runId": "run_abc123",
"status": "PENDING",
"model": "grok-imagine-video-1.5-i2v",
"estimatedCredits": 135,
"statusUrl": "/api/v1/runs/run_abc123"
}2 — Poll until COMPLETED (every 2-5 seconds)
curl https://app.scenetra.com/api/v1/runs/run_abc123 \
-H "Authorization: Bearer YOUR_API_KEY"Response
{
"runId": "run_abc123",
"status": "COMPLETED",
"progress": { "completed": 1, "total": 1 },
"outputs": [
{ "type": "video", "url": "https://images.scenetra.com/..." }
],
"error": null,
"estimatedCredits": 135
}Request fields
| Field | Type | Description | Default |
|---|---|---|---|
| prompt* | string | Describe the video: subject, motion, camera work, and audio direction. Native synchronized audio is generated in the same pass. | — |
| image_url* | string (https URL) | Starting frame. The video animates from this image. | — |
| resolution | string | Output resolution. 480p is the budget rung; 1080p is the premium rung.480p · 720p · 1080p | 720p |
| duration | string | Length of the generated video in seconds (1-15). Billed per second.1 · 2 · 3 · 4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 | 6 |
| callbackUrl | string (https URL) | Optional. When the run finishes (completed or failed), we POST the final run status to this URL — same shape as the poll response, plus model, params, and chargedCredits. Best-effort; polling remains the source of truth. | |
Unknown fields are rejected with a helpful 400. Media fields take https URLs. There is no provider parameter — every run is routed to the cheapest capable provider automatically.
Good to know
Billing
Runs bill your account credits at provider cost with 0% markup. The submit response includes the credit estimate; failed generations are refunded automatically. Add ?dryRun=1 to get the estimate without running anything.
Errors
Errors are always { "error": { "code", "message" } } with stable codes: unauthorized, model_not_available, invalid_parameter, insufficient_credits, rate_limited. A failed run is not an HTTP error — poll responses report status: "FAILED" with the reason.
What makes Grok Imagine Video 1.5 Image to Video different
Native audio, included in the price
Grok Imagine Video 1.5 generates synchronized audio with every clip in the same pass as the picture, and it's covered by the per-second rate rather than billed as an add-on. You can even direct the soundtrack from the prompt — name the ambience, effects, or mood you want and the model scores the scene to match what happens on screen.
Free-form durations from 1 to 15 seconds
Pick any length from 1 to 15 seconds instead of choosing between a couple of fixed presets. One- and two-second clips make cheap motion tests, six seconds is the sweet-spot default, and 15 seconds is enough for a full beat of action with sound — at 480p, a 15-second draft still costs well under $0.20.
Three resolutions, one cheap draft rung
480p runs at $0.012 per second, 720p at $0.0225, and 1080p at $0.04 — so a 5-second draft costs $0.06 and a 5-second full-HD final $0.20. A practical Scenetra workflow: iterate your prompt at 480p until motion and audio land, then flip the resolution select to 1080p for the final render. Same node, same wiring, and every generation is billed at provider cost with 0% markup.
Eight aspect ratios, or let the model decide
Grok Imagine Video 1.5 outputs 16:9, 4:3, 3:2, 1:1, 2:3, 3:4, and 9:16 — covering widescreen, square, and vertical social formats — or you can leave the ratio on Auto and let the model pick a frame that suits your prompt. That makes it easy to render the same idea as both a landscape hero clip and a 9:16 story cut.
Parameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Describe the video: subject, motion, camera work, and audio direction. Native synchronized audio is generated in the same pass. | — |
| Resolution | select | Output resolution. 480p is the budget rung; 1080p is the premium rung.480p · 720p · 1080p | 720p |
| Aspect Ratio | select | Aspect ratio of the generated video. Auto lets the model choose from your prompt.Auto · 16:9 · 4:3 · 3:2 · 1:1 · 2:3 · 3:4 · 9:16 | auto |
| Duration | select | Length of the generated video in seconds (1-15).1 · 2 · 3 · 4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 | 6 |
Pricing
From $0.012 per second
Pay per generation — only pay for what you use.
480p
$0.012/s
per second
720p
$0.0225/s
per second
1080p
$0.04/s
per second
What a video costs
| Duration | 480p | 720p | 1080p |
|---|---|---|---|
| 5 seconds | $0.06 | $0.113 | $0.20 |
| 10 seconds | $0.12 | $0.225 | $0.40 |
Use cases
Social content
- Sound-on hooks and openers
- Vertical 9:16 story clips
- Meme and reaction videos
- Daily posting at draft prices
Marketing & ads
- Concept ads with audio
- Product teaser clips
- A/B test multiple ad variants
- Seasonal campaign spots
Film & previz
- Establishing shots from a prompt
- Mood pieces with ambience
- Cheap 480p shot exploration
- Scene beats up to 15 seconds
Music & audio-led
- Visualizers scored by the model
- Ambient loops with sound
- Lyric-inspired scene clips
- Sound-design experiments
Frequently asked questions
What is Grok Imagine Video 1.5?+
It's xAI's latest image-to-video model: upload a still image, describe the motion in a prompt, and it generates a video of up to 15 seconds at 480p or 720p, with synchronized audio included when generated. xAI describes the 1.5 models as its best image-to-video yet, with improved motion and physics. On Scenetra it runs as a node in the visual workflow editor.
How do I use Grok Imagine to animate a photo?+
On Scenetra: drop a Grok Imagine Video v1.5 node onto your board, upload the photo, write a short prompt describing the motion you want — camera moves, actions, atmosphere — pick a duration and resolution, and generate. No X app or API key needed; it works in the browser.
Does Grok Imagine Video 1.5 generate audio?+
Yes — v1.5 generates synchronized audio with the video, and it's included in the per-second price rather than billed as an add-on. That covers ambient sound and effects that match the animated scene.
How should I prompt Grok Imagine for image to video?+
Describe the motion, not the image — the model already sees your upload. Name the action ('she turns toward the camera'), the camera move ('slow dolly in'), and the mood or sound you want. Short, concrete prompts tend to beat long scene descriptions, and Scenetra's board makes it cheap to iterate at 480p until the motion lands.
Related models
Seedance 2.5
ByteDance's Seedance 2.5 generates cinematic videos up to 30 seconds from text prompts with synchronized audio — voice, sound effects, and background music — built in by default. The newest Seedance generation, with mp4 or high-fidelity mov output for multi-round editing pipelines.
View model →Kling 3.0 Pro
Kuaishou's flagship video generation model delivering stunning visual fidelity. Kling 3.0 Pro supports text-to-video, image-to-video with start/end frames, element references for character consistency, and native audio generation.
View model →PixVerse v6
PixVerse's flagship video model generates 1-15 second clips from a text prompt with optional synchronized audio, multi-shot sequences, four resolutions from 360p to 1080p, and eight aspect ratios from vertical 9:16 to cinematic 21:9. Starting at $0.02 per second, it is one of the most affordable ways to produce finished AI video.
View model →Happy Horse 1.1
Alibaba's Happy Horse 1.1 generates 3-15 second videos with native audio straight from a text prompt. Dialogue, music, and sound effects are directed through the same prompt as the visuals, and nine aspect ratios from 9:21 to 21:9 cover everything from vertical stories to ultrawide cinematics at 480p, 720p, or 1080p.
View model →Start creating with Grok Imagine Video 1.5 Image to Video
Use Grok Imagine Video 1.5 Image to Video alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started Free