Image-to-Video with Native Audio with MiniMax H3
Animate any image into a 5-15 second video with MiniMax H3 (Hailuo 03). The clip starts exactly on your uploaded first frame, you can optionally pin a last frame to control where the motion lands, and stereo audio — dialogue, music, and SFX — is rendered natively in the same pass. Output from 480P up to 4K.
What makes MiniMax H3 Image-to-Video different
First frame in, optional last frame out
Your uploaded image becomes frame one, so the clip starts exactly on your composition, colors, and character. Add an optional last-frame image and H3 animates a coherent transition between the two — a controlled way to choreograph reveals, turns, and camera moves instead of hoping the motion drifts somewhere good.
Your still gets a soundtrack
H3 renders stereo audio natively in the same pass as the animation: dialogue for a portrait, ambience for a landscape, a music cue for a product shot. Describe the sound in the prompt alongside the motion and it arrives synchronized, at no extra cost on the per-second rate.
Draft cheap, deliver at up to 4K
768P is the cheapest rung at $0.04 per second — below even 480P — which makes it the natural resolution for testing whether an image animates well. When the motion lands, switch the same node to 2K ($0.065/s) or 4K ($0.16/s); both upscale a 768P base render for delivery-ready output.
Chain from any image model
On a Scenetra board, wire an image node — Nano Banana 2, Seedream, FLUX — straight into an H3 Image to Video node. Generate the perfect still, animate it with audio, and pass the result on to sequencing or editing nodes, all in one visual workflow with a live cost estimate on every node.
Playground
A timelapse of a flower blooming in a sunlit meadow, cinematic quality
Drop images or click to upload
Try MiniMax H3 Image-to-Video in Scenetra
Open PlaygroundParameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Describe the scene, action, camera movement, visual style, dialogue, music, and sound effects — H3 renders audio natively. | — |
| Image* | image upload | First-frame image. The video starts from this image. | — |
| Last Frame | image upload | Optional last-frame image. The video transitions from the first frame to this one. | — |
| Duration | select | Duration of the video in seconds (5-15).5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 | 5 |
| Resolution | select | 480P and 768P are native; 2K and 4K upscale a 768P base render.480P · 768P · 2K · 4K | 2K |
Pricing
From $0.04 per second, audio included
Pay per generation — only pay for what you use.
480P
$0.05/s
per second
768P
$0.04/s
per second
2K
$0.065/s
per second
4K
$0.16/s
per second
What a video costs
| Duration | 480P | 768P | 2K | 4K |
|---|---|---|---|---|
| 5 seconds | $0.25 | $0.20 | $0.33 | $0.80 |
| 10 seconds | $0.50 | $0.40 | $0.65 | $1.60 |
Use cases
Portraits & characters
- Talking portraits with real dialogue
- Character intros with music
- Animate generated art
- First-to-last-frame expression changes
Product & e-commerce
- Animate product photos with sound design
- Listing videos from stills
- Packshot-to-reveal transitions
- Seasonal campaign variants
Social content
- Bring photos to life with audio
- Sound-on hooks from a single still
- Vertical story clips
- Before/after transitions via last frame
Film & previz
- Animate concept art with temp sound
- Key-frame previz between two boards
- Establishing shots from stills
- Mood pieces with ambience
Related models
Happy Horse 1.1
Animate a first-frame image into a 3-15 second video with native audio using Alibaba's Happy Horse 1.1. The video starts exactly from your image and keeps its aspect ratio, while an optional prompt directs the motion, style, and soundtrack at 480p, 720p, or 1080p.
View model →Seedance 2.5 Image-to-Video
Animate any image into a cinematic video up to 30 seconds with Seedance 2.5 by ByteDance. Start from a first-frame image, optionally set a last frame to control the ending, and get synchronized audio by default. Output preserves your image's aspect ratio.
View model →Wan 3.0
Animate a still image into up to 30 seconds of video with Alibaba's Wan 3.0. The model uses your upload as the starting frame, follows your prompt for motion and camera work, and generates native synchronized audio in the same pass. Any duration from 2 to 30 seconds, at 480p, 720p, or 1080p — and the starting frame is free.
View model →Frequently asked questions
How do I animate an image with MiniMax H3?+
On Scenetra: drop a MiniMax H3 Image to Video node onto your board, upload the image, and write a prompt describing the motion, camera work, and any dialogue, music, or sound effects you want. Pick a duration (5-15 seconds) and resolution, then generate. The video starts exactly on your uploaded frame.
What does the last-frame option do?+
It pins where the clip ends. Upload a second image as the last frame and H3 animates a coherent transition from your first frame to it — useful for choreographed reveals, expression changes, camera moves between two compositions, or before/after transitions. It's optional; with only a first frame, the prompt drives the motion freely.
How much does MiniMax H3 image-to-video cost?+
Per second of output: $0.05 at 480P, $0.04 at 768P, $0.065 at 2K, and $0.16 at 4K, audio included. A 5-second 2K clip is about $0.33; a 10-second 4K clip is $1.60. Scenetra sells every generation at provider cost with 0% markup, pay-per-generation.
Does the animated video include audio?+
Yes. H3 renders stereo audio natively with the animation — dialogue, music, and sound effects described in your prompt, synchronized to the motion. There is no extra charge for audio; the per-second rate covers picture and sound together.
Which resolution should I pick?+
Start at 768P — it's the cheapest rung at $0.04 per second and a native render, so it's ideal for testing whether an image animates the way you want. 480P is also native; 2K and 4K upscale a 768P base render and are the delivery options once a take lands.
Can I generate the image and animate it in one workflow?+
Yes — that's what Scenetra's board is for. Wire an image model node (Nano Banana 2, Seedream 5, FLUX 2 Pro, and more) into the H3 Image to Video node, and the generated still flows straight in as the first frame. You can iterate on either stage independently and re-run only what changed.
Start creating with MiniMax H3 Image-to-Video
Use MiniMax H3 Image-to-Video alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started Free