Text-to-Video with Native Audio with MiniMax H3
MiniMax H3 is the newest generation of the Hailuo video lineage: 5-15 second clips with native stereo audio — dialogue, music, and sound effects rendered in the same pass as the picture. Generate at 480P or 768P natively, or step up to 2K and 4K upscaled output, in six aspect ratios from cinematic 21:9 to vertical 9:16.
What makes MiniMax H3 different
Audio is part of the generation, not an add-on
H3 renders stereo audio in the same pass as the picture: spoken dialogue, background music, and sound effects that match what's on screen. Write the sound into your prompt the way you'd write the visuals — a line of dialogue, a music cue, a door slam — and it arrives synchronized, with no separate audio pass and no surcharge on the per-second rate.
Four resolution rungs, from drafts to 4K
480P and 768P are native renders; 2K and 4K upscale a 768P base for delivery-ready output. Unusually, 768P is the cheapest rung at $0.04 per second — below even 480P — which makes it the natural default for iteration. When a take lands, flip the same node to 2K ($0.065/s) or 4K ($0.16/s) and re-run.
5 to 15 seconds, any whole second
Duration is a per-second choice from 5 to 15, not a couple of fixed presets. Five seconds is a cheap motion test; fifteen is enough for a full beat with dialogue and a music cue. Because billing is per second, a 7-second clip costs exactly 7 seconds — you never pay for a 10-second slot you didn't use.
Three modes on one board
Text to Video is one of three MiniMax H3 modes on Scenetra. Animate a first-frame image (optionally pinning the last frame) with Image to Video, or drive generations from up to 9 reference images, 3 reference clips, and 3 audio tracks with Reference to Video. All three run as nodes in the same visual workflow, so you can chain an image model straight into H3.
Playground
A timelapse of a flower blooming in a sunlit meadow, cinematic quality
Try MiniMax H3 in Scenetra
Open PlaygroundParameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Describe the scene, action, camera movement, visual style, dialogue, music, and sound effects — H3 renders audio natively. | — |
| Duration | select | Duration of the video in seconds (5-15).5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 | 5 |
| Resolution | select | 480P and 768P are native; 2K and 4K upscale a 768P base render.480P · 768P · 2K · 4K | 2K |
| Aspect Ratio | select | Aspect ratio of the generated video.21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16 | 16:9 |
Pricing
From $0.04 per second, audio included
Pay per generation — only pay for what you use.
480P
$0.05/s
per second
768P
$0.04/s
per second
2K
$0.065/s
per second
4K
$0.16/s
per second
What a video costs
| Duration | 480P | 768P | 2K | 4K |
|---|---|---|---|---|
| 5 seconds | $0.25 | $0.20 | $0.33 | $0.80 |
| 10 seconds | $0.50 | $0.40 | $0.65 | $1.60 |
Use cases
Film & storytelling
- Dialogue scenes with synced voices
- Establishing shots in 21:9
- Mood pieces with score and ambience
- Previz with temp sound built in
Marketing & ads
- Product spots with music and VO
- UGC-style ads with natural dialogue
- Brand stings with sound design
- A/B variants at draft resolution
Social content
- Vertical 9:16 clips with audio hooks
- Talking-character shorts
- Sound-on meme clips
- Music-driven transitions
Music & performance
- Performance clips with generated music
- Rhythm-matched visuals
- Lyric-video backgrounds
- Ambient loops with soundscapes
Related models
Happy Horse 1.1
Alibaba's Happy Horse 1.1 generates 3-15 second videos with native audio straight from a text prompt. Dialogue, music, and sound effects are directed through the same prompt as the visuals, and nine aspect ratios from 9:21 to 21:9 cover everything from vertical stories to ultrawide cinematics at 480p, 720p, or 1080p.
View model →Seedance 2.5
ByteDance's Seedance 2.5 generates cinematic videos up to 30 seconds from text prompts with synchronized audio — voice, sound effects, and background music — built in by default. The newest Seedance generation, with mp4 or high-fidelity mov output for multi-round editing pipelines.
View model →Kling 3.0 Pro
Kuaishou's flagship video generation model delivering stunning visual fidelity. Kling 3.0 Pro supports text-to-video, image-to-video with start/end frames, element references for character consistency, and native audio generation.
View model →Grok Imagine Video 1.5
xAI's Grok Imagine Video 1.5 turns a text prompt into a video with synchronized audio generated in the same pass — no separate sound step and no audio surcharge. Choose any duration from 1 to 15 seconds, three resolutions from 480p to 1080p, and eight aspect ratios from widescreen 16:9 to vertical 9:16.
View model →Frequently asked questions
What is MiniMax H3?+
MiniMax H3 (Hailuo 03) is the latest video model in MiniMax's Hailuo lineage. It generates 5-15 second videos from a text prompt with native stereo audio — dialogue, music, and sound effects rendered together with the picture — at resolutions from 480P up to 4K. On Scenetra it runs as a node in the visual workflow editor, alongside Image to Video and Reference to Video modes.
How much does MiniMax H3 cost?+
Per second of output: $0.05 at 480P, $0.04 at 768P, $0.065 at 2K, and $0.16 at 4K, with audio included. A 5-second 2K clip is about $0.33 and a 10-second one about $0.65. Scenetra bills every generation at provider cost with 0% markup, pay-per-generation with no subscription gate.
Does MiniMax H3 generate audio?+
Yes — natively. H3 renders stereo audio in the same pass as the video: spoken dialogue, background music, and sound effects that stay synchronized with the action. There's no separate audio model, no extra toggle, and no surcharge — the per-second price covers picture and sound.
What resolutions does MiniMax H3 support?+
Four rungs: 480P and 768P are native renders, while 2K and 4K upscale a 768P base render for higher-resolution delivery. Pricing-wise 768P is actually the cheapest option at $0.04 per second, so it's the sweet spot for iterating; 2K is the default on the node.
How long can a MiniMax H3 video be?+
Any whole-second duration from 5 to 15 seconds per generation. Billing is per second, so a 7-second clip costs exactly 7 seconds' worth. If you need longer sequences, generate multiple clips on a Scenetra board and chain them with the sequence tools.
How should I prompt MiniMax H3 for audio?+
Treat sound as part of the scene description. Name the dialogue in quotes, describe the music ('slow piano, melancholic'), and call out effects ('rain on a tin roof, distant thunder'). Because H3 renders audio and video together, the model times sound to the on-screen action — you direct it the same way you direct the camera.
Is MiniMax H3 free to try?+
You can try it free through Scenetra: new accounts get free welcome credits and a 7-day trial that covers your first generations. After that it's pay-per-generation at provider cost — a 5-second 768P draft is about $0.20 — with nothing to cancel.
Start creating with MiniMax H3
Use MiniMax H3 alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started Free