MiniMax H3 prompt guide: control every reference input
Learn MiniMax H3 prompting for image, video, audio, and text references. Bind each asset to a role, prevent conflicts, and direct a timed AI video shot.

The useful idea behind a MiniMax H3 prompt is not “write more detail.” It is “give every input one job.” H3 can take text, images, video, and audio as a unified reference set, so a creator can draw product identity from a still, movement from a clip, pacing from audio, and shot direction from the prompt. That flexibility only helps when the model can tell which source controls which part of the result.
This MiniMax H3 prompt guide turns that principle into a repeatable format. You will learn how to bind each asset to a role, prevent references from contradicting one another, write preservation rules, and divide a 4–15 second generation into timed beats. Six templates at the end cover product video, character motion, dialogue, video editing, UI motion, and first-to-last-frame transitions.
Note
This is a documentation-led prompt guide, not a report of H3 generations tested by OmniArt. It is based on MiniMax's current H3 generation guide and official API reference. H3 is not currently confirmed in OmniArt's source-of-truth model catalog as of July 31, 2026. Interface labels and reference syntax may also differ between MiniMax's own surfaces and downstream hosts.
MiniMax H3, Hailuo 3.0, or Hailuo 03?
MiniMax's current documentation calls the model MiniMax H3 and gives the official API model identifier as MiniMax-H3. You may also encounter Hailuo 3.0, Hailuo 03, or Hailuo-3.0 in launch coverage and third-party catalogs because Hailuo is MiniMax's video product brand.
Use the label required by the surface in front of you. For the official MiniMax V2 API documented at publication time, that means MiniMax-H3; do not assume a marketplace alias is also a valid official API identifier. This guide uses MiniMax H3 for the model and reference generation for the mode that combines image, video, and audio inputs. “Omni reference” is a useful description of that mixed-media workflow, but the official API calls it reference-to-video.
Know the three H3 generation modes
Choosing the right mode comes before writing the prompt.
| Mode | Inputs | Use it when |
|---|---|---|
| Text-to-video | Required text prompt | The scene can be created from description alone |
| First/last-frame image-to-video | Required prompt plus a first frame, a last frame, or both | An exact opening or ending composition matters |
| Reference generation | Required prompt plus reference images, videos, or audio | You need identity, design, movement, camera, voice, rhythm, or style from uploaded assets |
The current official specification lists 2K output and integer durations from 4 to 15 seconds. Every request requires a non-empty text prompt, with a limit of 7,000 characters. A reference request can include up to nine images, three videos, and three audio clips, with no more than 12 assets total. Reference videos are limited to 15 seconds in aggregate, as are reference audio clips. An audio reference cannot be submitted alone; it must accompany at least one image or video.
Text-to-video requires a concrete aspect ratio. Reference generation can use a concrete ratio or let H3 choose adaptively, while first/last-frame generation follows the uploaded image. The documented ratio choices are 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16.
There is one hard workflow boundary: first/last-frame mode and reference generation cannot be mixed in the same API request. If exact boundary frames matter, use first/last-frame mode. If identity, motion, style, or audio transfer matters more, use reference generation and describe the opening composition in text.
Build an H3 prompt as a role map
H3's API receives media as typed items with roles such as reference_image, reference_video, and reference_audio. Those technical roles identify the media type; your prompt should go one level further and state the creative responsibility of each asset.
Use this order:
DELIVERABLE
What the final clip is for, its duration, and its format.
ROLE MAP
Reference image 1 = subject identity and immutable appearance.
Reference image 2 = environment or visual style only.
Reference video 1 = body motion and timing only.
Reference audio 1 = rhythm and edit timing only.
SHOT
Subject, action, environment, lighting, and camera direction.
TIMELINE
Time-coded beats that add up to the selected duration.
PRESERVE
Features that must remain unchanged from named references.
EXCLUDE
A short list of the most damaging unwanted changes.
The numbered labels above are semantic labels for your brief, not a promise that every H3 interface exposes literal Reference image 1 handles. In the official API, the files are separate items in the content array. In a visual interface, use whatever attachment names or mention syntax that surface provides, then keep the same one-asset-one-job logic.
Give each asset one primary job
A reference contains more information than you usually want to transfer. A dance clip contains a performer, clothes, location, camera, lighting, motion, and perhaps music. If you ask H3 to “follow the video,” all of those features compete with your character image.
Bind narrowly instead:
- Image reference: identity, face, wardrobe, product geometry, logo, palette, or environment design.
- Video reference: body motion, camera path, action timing, physical interaction, or editing rhythm.
- Audio reference: voice character, cadence, music rhythm, ambience, or an event cue.
- Text prompt: the final scene, transfer rules, timeline, priorities, and preservation instructions.
When one asset must do two jobs, name both and make the boundary explicit: “Reference video 1 controls camera path and action timing; ignore its performer, wardrobe, location, color grade, and audio.”
Prevent references from fighting each other
Reference conflict is usually an instruction problem before it is a model problem. Two images may show different jackets, a motion clip may use a different body type, or an audio beat may imply cuts that clash with a requested unbroken camera move.
Use a written priority order whenever inputs overlap:
Priority order:
1. Reference image 1 controls character identity and wardrobe.
2. Reference video 1 controls movement and timing only.
3. Reference image 2 controls lighting palette and set design only.
4. The text prompt controls camera framing and final composition.
Then resolve contradictions rather than hoping H3 will infer your preference. If the identity image has loose hair but the motion video has a hat, state whether the final character keeps the hair or wears the hat. If a product reference shows a readable label but the style image is painterly, say that stylization applies to the environment while packaging remains photorealistic.
Three habits reduce conflict:
- Use the smallest useful reference set. More inputs create more possible disagreements. Add a file only when it contributes information the prompt cannot express reliably.
- Separate content from style. Name which reference owns the subject and which owns the treatment.
- Choose one camera authority. Either transfer the camera from a video or direct it in text. If you need both, state that the video supplies timing while text overrides framing.
For a broader symptom-by-symptom workflow, keep 7 AI video prompt fixes beside this guide while you iterate.
Write preservation instructions that can be checked
“Keep it consistent” is too vague. A preservation block should name visible properties you can inspect in any frame.
For a character:
Preserve from reference image 1 throughout every frame: facial structure,
eye color, hairstyle and length, jacket cut, jacket color, and body proportions.
Do not replace the performer with the person from reference video 1.
For a product:
Preserve from reference image 1: bottle silhouette, cap shape, label placement,
navy wordmark, coral seal, material finish, and relative proportions.
No duplicate product, redesigned packaging, extra text, or changing logo.
Positive anchors should come first; exclusions are a guardrail, not the whole prompt. Keep the negative list short and specific. Ten vague prohibitions can dilute the four invariants that actually decide whether a take is usable.
If the same character will appear across several separately generated clips, keep the identity and wardrobe block identical in every prompt. The consistent-character workflow explains how to build that reusable reference packet across a longer sequence.
Direct time, not just content
H3 allows clips from 4 to 15 seconds, but a long paragraph does not tell the model when each event happens. A timeline turns the prompt into a compact shot plan.
For a 10-second clip:
0–2 seconds: locked medium-wide establishing shot; subject holds still.
2–6 seconds: subject performs the referenced action at the source tempo.
6–9 seconds: camera makes one slow 20-degree arc to the right.
9–10 seconds: subject settles; hold a clean final composition for the edit.
Keep each beat achievable. One main action plus one camera move is usually clearer than a chain of unrelated transformations. Let the time ranges add up to the selected duration, place important product or facial details in a slower beat, and reserve the final half-second or second for a stable edit point.
Audio can share the same clock:
At 2.0 seconds, the first downbeat starts the hand movement.
At 6.0 seconds, the bass hit motivates the camera arc.
From 9.0 seconds, let the music tail continue under the held final frame.
For more lens, light, and camera vocabulary, see the cinematic AI video prompt guide.
Six MiniMax H3 prompt templates
Replace the bracketed details, attach only the listed assets, and adapt the labels to the interface you use. These templates are structured starting points based on documented H3 input modes; they are not claimed benchmark winners or guaranteed outputs.
Template 1: Product reveal with motion and music references
- Mode: Reference generation
- Assets: One clean product image, one camera-motion video, optional music reference
Create a 10-second 9:16 product reveal for a paid social placement.
Role map:
- Reference image 1 controls the product's exact design, proportions, packaging,
colors, materials, label placement, and wordmark.
- Reference video 1 controls camera path and acceleration only. Ignore its subject,
location, lighting, color grade, and audio.
- Reference audio 1 controls beat and edit timing only.
Scene: The product stands centered on a warm-violet studio plinth. Soft lilac key
light from camera left, restrained coral rim light, subtle atmospheric haze.
Timeline:
- 0–2 seconds: static wide reveal on the first soft beat.
- 2–7 seconds: transfer the smooth push-and-arc movement from reference video 1.
- 7–9 seconds: one narrow highlight travels across the product surface.
- 9–10 seconds: camera settles; hold the label front-facing and readable.
Preserve reference image 1 exactly: silhouette, cap, material finish, label layout,
brand colors, and wordmark. Keep one product only. No redesigned packaging,
duplicate objects, extra text, warped label, or abrupt camera shake.
For the still-image preparation that comes before this prompt, use the photo-to-product-video workflow.
Template 2: Character motion transfer without identity drift
- Mode: Reference generation
- Assets: One character image, one motion video, optional environment image
Create an 8-second 16:9 cinematic character performance.
Priority order:
1. Reference image 1 controls face, hair, body proportions, and wardrobe.
2. Reference video 1 controls body choreography and action timing only.
3. Reference image 2 controls the environment palette and architecture only.
4. This prompt controls framing and lighting.
The character from reference image 1 performs the complete movement from reference
video 1 in a moonlit station based on reference image 2. Medium full shot, camera
locked at chest height, soft directional light, realistic weight and foot contact.
Timeline:
- 0–1 seconds: character holds the starting pose.
- 1–7 seconds: perform the reference choreography once at its original tempo.
- 7–8 seconds: settle naturally and look toward camera left.
Preserve facial structure, hairstyle, coat shape, coat color, boots, and body
proportions from reference image 1 in every frame. Do not inherit the motion video's
performer, face, clothing, background, camera movement, or audio. No extra limbs,
sliding feet, costume changes, or cuts.
Template 3: Dialogue performance with a voice reference
- Mode: Reference generation
- Assets: One character image, one authorized voice or dialogue audio clip
Warning
Only upload a voice you own or have explicit permission to use. A model accepting reference audio does not give you rights to imitate another person's voice.
Create a 9-second 16:9 single-character dialogue shot.
Role map:
- Reference image 1 controls the speaker's appearance, wardrobe, and room design.
- Reference audio 1 controls spoken timing, cadence, emotional progression, and pauses.
Shot: Medium close-up, eye-level, 50 mm cinematic framing. The speaker begins calm,
briefly smiles after the central pause, then finishes with quiet confidence. Natural
blinks and restrained hand movement. Camera remains static; soft room tone underneath.
Timeline:
- 0–1 seconds: silent eye contact and a small inhale.
- 1–8 seconds: performance follows reference audio 1 exactly in timing and pauses.
- 8–9 seconds: mouth closes, expression settles, hold for the edit.
Preserve face, hairstyle, skin tone, jacket, background layout, and lighting direction
from reference image 1. Keep one speaker. No camera move, cutaway, background speech,
new words, exaggerated gestures, or wardrobe changes.
Template 4: Replace a source-video environment
- Mode: Reference generation
- Assets: One source video, one environment image, optional subject image
Create a 12-second 16:9 environmental replacement based on the source clip.
Role map:
- Reference video 1 controls shot length, subject action, physical timing, camera path,
framing progression, and interaction with the ground.
- Reference image 1 controls the replacement environment, architecture, palette,
weather, and lighting mood.
- Reference image 2 controls the subject's face and wardrobe, if supplied.
Replace the source video's location with the environment from reference image 1.
Preserve the original action and camera movement continuously. Match subject lighting,
contact shadows, reflections, and atmospheric perspective to the new environment.
Timeline follows reference video 1. Maintain one continuous shot with the original
action beats and no added event.
Preserve the source subject's position, scale, motion, and ground contact. Preserve
identity and wardrobe from reference image 2 when present. Do not copy the original
background, signage, bystanders, color grade, or source audio. No cuts, teleporting,
floating feet, or changing architecture.
Template 5: Branded UI motion synchronized to a beat
- Mode: Reference generation
- Assets: One UI layout image, one motion-style video, optional audio reference
Create a 7-second 1:1 UI motion-design clip for a product announcement.
Priority order:
1. Reference image 1 controls exact layout, hierarchy, component positions, text,
logo, colors, corner shapes, and typography appearance.
2. Reference video 1 controls transition character and easing only.
3. Reference audio 1 controls the timing of three motion beats only.
Animate the interface from reference image 1 without redesigning it. Begin with the
main panel at rest. Use the restrained slide-and-scale transition quality from
reference video 1. Keep the camera orthographic and the background static.
Timeline:
- 0–1 seconds: complete layout at rest.
- 1–3 seconds: cards enter in sequence on beat one.
- 3–5 seconds: primary control changes state on beat two.
- 5–6 seconds: one subtle emphasis pulse on beat three.
- 6–7 seconds: complete interface holds sharp and readable.
Preserve every word, logo shape, component proportion, spacing relationship, and brand
color from reference image 1. No invented labels, misspelled text, extra panels,
perspective tilt, camera movement, glow overload, or elastic distortion.
Template 6: Exact first-to-last-frame transformation
- Mode: First/last-frame image-to-video
- Assets: One first-frame image and one last-frame image; no reference media
Create a 10-second transformation from the supplied first frame to the supplied last
frame. Preserve the subject's identity and the camera's fixed position throughout.
Timeline:
- 0–2 seconds: hold the first-frame composition; only subtle ambient movement.
- 2–7 seconds: the scene transforms progressively from the center outward. Materials
change continuously with believable physical contact and no hard cut.
- 7–9 seconds: remaining details resolve into the last-frame design.
- 9–10 seconds: arrive exactly at the supplied last frame and hold it cleanly.
One continuous locked shot. Preserve subject scale, face, silhouette, horizon, lens,
and framing during the transition. No reference-style transfer, new characters,
camera movement, jump cut, flicker, or overshoot beyond the final composition.
Do not attach the other templates' reference images, videos, or audio to this request. The official API treats first/last-frame generation and reference generation as mutually exclusive modes.
A practical iteration order
Do not change identity, motion, camera, style, and audio instructions after one disappointing render. That makes it impossible to learn which instruction helped.
Iterate in three passes:
- Lock roles and invariants. Use the smallest reference set and verify that the subject or product stays recognizable.
- Lock action and timeline. Add the motion source or timed beats, keeping style simple.
- Add treatment. Introduce environment, lighting, grade, audio cues, and finishing details after the first two layers are stable.
When a result drifts, simplify before adding more negative language. Remove the least important reference, shorten the timeline, or give one source a narrower job. Compare several variants against the same checklist: identity, geometry, motion, camera, audio timing, and final-frame stability.
Getting started from OmniArt
MiniMax H3 is not currently confirmed as selectable inside OmniArt, so this article does not link to an H3 creation route or claim an OmniArt H3 workflow. Check the current OmniArt video model lineup for the models available today.
The role-map method still transfers well to supported reference-driven video models: prepare one identity image, one motion source, and a short preservation block before opening the workspace. Keep the prompt model-neutral until the creative brief is stable, then adapt attachment labels and feature-specific syntax to the model you actually select. That gives you a reusable direction sheet now and a clean H3 prompt if the model joins OmniArt later.
For official API rates, reference-video billing, and worked cost examples, use the MiniMax H3 pricing and specs guide. It uses MiniMax's own pricing page rather than a reseller rate.
Ready to Create?
Start generating amazing content with AI