Reference-to-Video with MiniMax H3
Direct MiniMax H3 (Hailuo 03) with the material you already have: up to 9 reference images, 3 reference video clips, and 3 audio tracks in a single generation. Keep subjects and products consistent across shots, borrow motion and pacing from reference footage, sync action to a music track, and get native audio in the output — 5-15 seconds, from 480P up to 4K.
What makes MiniMax H3 Reference-to-Video different
Up to 15 references in one generation
Load up to 9 images, 3 video clips, and 3 audio tracks into a single H3 generation. Images lock in subjects, products, style, or composition; video clips contribute motion, pacing, or editing rhythm; audio tracks drive timing, voice, or music. Cite each one in your prompt by slot — 'Image 1 wears the jacket from Image 2, moving like Video 1' — and the model composes them into one coherent clip.
Consistent characters and products across shots
Reference images are how you keep the same face, outfit, or product across a whole campaign. Reuse the identical reference set across generations on a Scenetra board, vary only the prompt per shot, and H3 holds the subject steady — the practical route to multi-shot stories and product families that actually match.
Cut to the beat with audio references
Attach up to 3 audio tracks (2-15 seconds each, 15 seconds combined) and H3 syncs motion, voice, or action to them — pair a music track with your subject references and the clip moves on the beat. Audio references need at least one image or video reference alongside, and H3 still renders its own native audio in the output.
Auto aspect ratio follows your material
Leave Aspect Ratio on Auto and the output follows your reference media's framing — vertical references produce vertical video without manual matching. When you need a specific delivery format, override it with any of six ratios from 21:9 to 9:16, and choose 480P/768P native or 2K/4K upscaled output at 5-15 seconds. Rates match the other H3 modes: 768P is the cheapest rung at $0.04 per second, 2K is $0.065, and 4K tops out at $0.16.
Playground
A timelapse of a flower blooming in a sunlit meadow, cinematic quality
Drop images or click to upload
Try MiniMax H3 Reference-to-Video in Scenetra
Open PlaygroundParameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Describe the result and cite reference inputs in order: Image 1, Video 1, Audio 1, and so on. Audio (dialogue, music, SFX) is rendered natively. | — |
| Reference Images | image upload | Up to 9 reference images for subjects, products, style, or composition. Cite in the prompt as Image 1, Image 2, and so on. | — |
| Reference Videos | video upload | Up to 3 reference clips (2-15s each, combined 15s or less) for motion, pacing, or editing rhythm. Cite as Video 1, Video 2, and so on. | — |
| Reference Audio | audio upload | Up to 3 audio tracks (2-15s each, combined 15s or less) to sync motion, voice, or music. Needs at least one image or video reference alongside. Cite as Audio 1, Audio 2, and so on. | — |
| Duration | select | Duration of the video in seconds (5-15).5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 | 5 |
| Resolution | select | 480P and 768P are native; 2K and 4K upscale a 768P base render.480P · 768P · 2K · 4K | 2K |
| Aspect Ratio | select | Aspect ratio. 'Auto' follows the reference media.Auto · 21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16 | Auto |
Pricing
From $0.04 per second, audio included
Pay per generation — only pay for what you use.
480P
$0.05/s
per second
768P
$0.04/s
per second
2K
$0.065/s
per second
4K
$0.16/s
per second
What a video costs
| Duration | 480P | 768P | 2K | 4K |
|---|---|---|---|---|
| 5 seconds | $0.25 | $0.20 | $0.33 | $0.80 |
| 10 seconds | $0.50 | $0.40 | $0.65 | $1.60 |
Use cases
Brand & product campaigns
- Product families with matching looks
- Same product across scene variations
- Style-locked campaign series
- Logo and packshot continuity
Characters & storytelling
- Recurring characters across shots
- Multi-shot story sequences
- Wardrobe and identity consistency
- Casting from a single portrait
Music & rhythm
- Clips cut to a music track
- Dance motion from reference footage
- Beat-synced product reveals
- Voice-driven character performance
Motion transfer
- Borrow pacing from reference clips
- Re-stage motion with new subjects
- Match an editing rhythm
- Camera-move replication
Related models
Happy Horse 1.1
Generate 3-15 second videos from up to 9 reference images with Alibaba's Happy Horse 1.1. Cite each reference positionally in the prompt — [Image 1], [Image 2] — to lock subjects, products, or styles into the scene, with native audio and nine aspect ratios from 9:21 to 21:9.
View model →Seedance 2.5 Reference-to-Video
Generate videos from up to 30 reference images, 10 reference videos, and 10 audio tracks with Seedance 2.5 by ByteDance. Cite references directly in your prompt as @Image1, @Video1, @Audio1. Unique to 2.5: audio-only referencing — a single music or voice track can drive visual pacing, beat matching, and lip-sync.
View model →Wan 3.0
Generate video from up to 10 reference images with Alibaba's Wan 3.0. The model carries the people, objects, and styles from your references into a continuous take of 2 to 30 seconds, with native synchronized audio generated in the same pass. Output at 480p, 720p, or 1080p — reference images are free, and audio adds nothing to the price.
View model →Frequently asked questions
What is MiniMax H3 Reference to Video?+
It's the H3 mode that generates video from reference material instead of just a prompt: up to 9 reference images, 3 reference video clips, and 3 audio tracks in one generation. Images anchor subjects and style, videos contribute motion and pacing, and audio drives timing — while H3 still renders native audio in the output. On Scenetra it runs as a node in the visual workflow editor.
How do I reference the uploads in my prompt?+
By slot, in order: Image 1, Image 2, Video 1, Audio 1, and so on. Write the prompt as direction — 'the woman from Image 1 walks through the cafe from Image 3, moving to the rhythm of Audio 1' — so the model knows which reference plays which role in the scene.
What are the limits on reference videos and audio?+
Up to 3 reference video clips, each 2-15 seconds, with a combined length of 15 seconds or less — and the same limits for audio: up to 3 tracks, 2-15 seconds each, 15 seconds combined. Audio references also need at least one image or video reference alongside them in the same generation.
How does pricing work for Reference to Video?+
Per second of output, the same rates as the other H3 modes: $0.05 at 480P, $0.04 at 768P, $0.065 at 2K, and $0.16 at 4K — audio included. A 5-second 2K clip is about $0.33 and a 10-second one about $0.65. Scenetra bills every generation at provider cost with 0% markup, pay-per-generation with no subscription gate.
Can I keep the same character across multiple videos?+
Yes — that's the core use case. Put the character's reference images on a Scenetra board, wire them into several H3 Reference to Video nodes, and change only the prompt per shot. Because every generation sees the same references, the face, outfit, and styling stay consistent across the whole sequence.
What resolutions and durations are available?+
Any whole-second duration from 5 to 15 seconds, at 480P or 768P native, or 2K and 4K upscaled from a 768P base render. Aspect ratio can be set explicitly (21:9 through 9:16) or left on Auto to follow the framing of your reference media.
Start creating with MiniMax H3 Reference-to-Video
Use MiniMax H3 Reference-to-Video alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started Free