Text to VideoImage to VideoReference to Video

Reference-to-Video with MiniMax H3

Direct MiniMax H3 (Hailuo 03) with the material you already have: up to 9 reference images, 3 reference video clips, and 3 audio tracks in a single generation. Keep subjects and products consistent across shots, borrow motion and pacing from reference footage, sync action to a music track, and get native audio in the output — 5-15 seconds, from 480P up to 4K.

What makes MiniMax H3 Reference-to-Video different

Up to 15 references in one generation

Load up to 9 images, 3 video clips, and 3 audio tracks into a single H3 generation. Images lock in subjects, products, style, or composition; video clips contribute motion, pacing, or editing rhythm; audio tracks drive timing, voice, or music. Cite each one in your prompt by slot — 'Image 1 wears the jacket from Image 2, moving like Video 1' — and the model composes them into one coherent clip.

Consistent characters and products across shots

Reference images are how you keep the same face, outfit, or product across a whole campaign. Reuse the identical reference set across generations on a Scenetra board, vary only the prompt per shot, and H3 holds the subject steady — the practical route to multi-shot stories and product families that actually match.

Cut to the beat with audio references

Attach up to 3 audio tracks (2-15 seconds each, 15 seconds combined) and H3 syncs motion, voice, or action to them — pair a music track with your subject references and the clip moves on the beat. Audio references need at least one image or video reference alongside, and H3 still renders its own native audio in the output.

Auto aspect ratio follows your material

Leave Aspect Ratio on Auto and the output follows your reference media's framing — vertical references produce vertical video without manual matching. When you need a specific delivery format, override it with any of six ratios from 21:9 to 9:16, and choose 480P/768P native or 2K/4K upscaled output at 5-15 seconds. Rates match the other H3 modes: 768P is the cheapest rung at $0.04 per second, 2K is $0.065, and 4K tops out at $0.16.

Playground

MiniMax H3 Reference-to-Video
playground

A timelapse of a flower blooming in a sunlit meadow, cinematic quality

Drop images or click to upload

5
2K
Auto
Generate
Resultpreview
Sample output from MiniMax H3 Reference-to-Video

Try MiniMax H3 Reference-to-Video in Scenetra

Open Playground

Parameters

ParameterTypeDescriptionDefault
Prompt*textDescribe the result and cite reference inputs in order: Image 1, Video 1, Audio 1, and so on. Audio (dialogue, music, SFX) is rendered natively.
Reference Imagesimage uploadUp to 9 reference images for subjects, products, style, or composition. Cite in the prompt as Image 1, Image 2, and so on.
Reference Videosvideo uploadUp to 3 reference clips (2-15s each, combined 15s or less) for motion, pacing, or editing rhythm. Cite as Video 1, Video 2, and so on.
Reference Audioaudio uploadUp to 3 audio tracks (2-15s each, combined 15s or less) to sync motion, voice, or music. Needs at least one image or video reference alongside. Cite as Audio 1, Audio 2, and so on.
DurationselectDuration of the video in seconds (5-15).5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 155
Resolutionselect480P and 768P are native; 2K and 4K upscale a 768P base render.480P · 768P · 2K · 4K2K
Aspect RatioselectAspect ratio. 'Auto' follows the reference media.Auto · 21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16Auto

Pricing

From $0.04 per second, audio included

Pay per generation — only pay for what you use.

480P

$0.05/s

per second

768P

$0.04/s

per second

2K

$0.065/s

per second

4K

$0.16/s

per second

What a video costs

Duration480P768P2K4K
5 seconds$0.25$0.20$0.33$0.80
10 seconds$0.50$0.40$0.65$1.60

Use cases

Brand & product campaigns

  • Product families with matching looks
  • Same product across scene variations
  • Style-locked campaign series
  • Logo and packshot continuity

Characters & storytelling

  • Recurring characters across shots
  • Multi-shot story sequences
  • Wardrobe and identity consistency
  • Casting from a single portrait

Music & rhythm

  • Clips cut to a music track
  • Dance motion from reference footage
  • Beat-synced product reveals
  • Voice-driven character performance

Motion transfer

  • Borrow pacing from reference clips
  • Re-stage motion with new subjects
  • Match an editing rhythm
  • Camera-move replication

Related models

Frequently asked questions

What is MiniMax H3 Reference to Video?+

It's the H3 mode that generates video from reference material instead of just a prompt: up to 9 reference images, 3 reference video clips, and 3 audio tracks in one generation. Images anchor subjects and style, videos contribute motion and pacing, and audio drives timing — while H3 still renders native audio in the output. On Scenetra it runs as a node in the visual workflow editor.

How do I reference the uploads in my prompt?+

By slot, in order: Image 1, Image 2, Video 1, Audio 1, and so on. Write the prompt as direction — 'the woman from Image 1 walks through the cafe from Image 3, moving to the rhythm of Audio 1' — so the model knows which reference plays which role in the scene.

What are the limits on reference videos and audio?+

Up to 3 reference video clips, each 2-15 seconds, with a combined length of 15 seconds or less — and the same limits for audio: up to 3 tracks, 2-15 seconds each, 15 seconds combined. Audio references also need at least one image or video reference alongside them in the same generation.

How does pricing work for Reference to Video?+

Per second of output, the same rates as the other H3 modes: $0.05 at 480P, $0.04 at 768P, $0.065 at 2K, and $0.16 at 4K — audio included. A 5-second 2K clip is about $0.33 and a 10-second one about $0.65. Scenetra bills every generation at provider cost with 0% markup, pay-per-generation with no subscription gate.

Can I keep the same character across multiple videos?+

Yes — that's the core use case. Put the character's reference images on a Scenetra board, wire them into several H3 Reference to Video nodes, and change only the prompt per shot. Because every generation sees the same references, the face, outfit, and styling stay consistent across the whole sequence.

What resolutions and durations are available?+

Any whole-second duration from 5 to 15 seconds, at 480P or 768P native, or 2K and 4K upscaled from a 768P base render. Aspect ratio can be set explicitly (21:9 through 9:16) or left on Auto to follow the framing of your reference media.

Start creating with MiniMax H3 Reference-to-Video

Use MiniMax H3 Reference-to-Video alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.

Get Started Free