guideTutorials & how-to guides14 min read

MiniMax H3 native audio guide: dialogue, music and sync

Learn how to prompt MiniMax H3 native stereo audio for dialogue, ambience, music, voice references, and sync, with a repeatable production test plan today.

OmniArt Team
MiniMax H3 native audio guide: dialogue, music and sync

MiniMax H3 native audio turns the soundtrack into part of the generation brief instead of a file you automatically add after the picture is finished. MiniMax says H3 jointly understands text, images, video, and audio, then generates videos with native stereo sound. Its training scope also treats voice, sound effects, and music as one audio domain rather than separate tools.

That is a meaningful workflow change, but it is not a promise that every line will be perfectly spoken, every footstep will land on the right frame, or every stereo mix will be ready to publish. This guide separates MiniMax's documented H3 capabilities from the production tests creators should run. It gives you a practical sound-brief structure, an audio-reference workflow, five prompt templates, and a repeatable evaluation scorecard without presenting official demos as independent results.

What MiniMax officially documents about H3 audio

MiniMax describes H3 as a general-purpose multimodal video model, not a silent video generator with a separate sound pass. The company's launch materials say the model generates native stereo audio and jointly models voice, sound effects, and music. MiniMax also demonstrates a prompt that combines camera motion from a video, a character from an image, and vocals matched to an audio reference.

The official H3 video generation guide defines the current API boundaries:

Documented capabilityWhat it means for a creator
Native stereo audioH3 can return a two-channel soundtrack with the generated video
Audio as multimodal contextAn audio clip can guide voice, motion, editing rhythm, or another relationship described in the prompt
Up to three audio referencesEach clip can be 2–15 seconds, with 15 seconds maximum across all audio references
Mixed referencesAudio can be combined with images and videos in one reference-to-video request
WAV and MP3 inputReference audio must use one of the documented input formats
4–15 second outputH3 is designed for short clips, not a complete long-form soundtrack in one generation

There are three important constraints. An audio reference cannot be submitted by itself; MiniMax requires at least one image or video alongside it. Audio files are limited to 15 MB each, and the complete mixed-reference request is limited to 12 assets. Reference mode also cannot be combined with first-frame or last-frame mode in the same request, so decide whether exact boundary frames or broader multimodal conditioning matters more for the shot.

Note

Stereo is not the same as spatial audio. MiniMax documents native stereo output, but its current public materials do not promise surround, ambisonic, object-based audio, or a particular level of positional accuracy. Evaluate left-right placement with headphones before describing a result as spatial.

Build the sound brief before writing the prompt

The most useful H3 prompt is not a long list of sounds. It is a compact sound plan that tells the model what should lead, what should support it, and what must happen at a visible moment. Write the plan in five passes.

1. Lock the visual action and timing

Audio synchronization depends on visible events. Start with the shot, subject action, camera movement, and sequence of beats. “A spoon hits the saucer as the camera settles” gives the model a shared audiovisual event. “Add a cup sound” does not say when the event happens or what should cause it.

For a 10-second clip, a simple beat map might be:

  • 0–3 seconds: establish the room and subject
  • 3–7 seconds: dialogue or primary action
  • 7–10 seconds: product reveal and final sound cue

Treat those timings as direction, not frame-accurate edit markers. A production test should measure how closely the output follows them.

2. Name one primary audio layer

Choose the sound the audience must notice first: dialogue, a product interaction, a musical hook, or a piece of environmental sound. Give it the clearest wording and protect space around it.

If dialogue is primary, write the exact line in quotation marks and specify one performance direction. If a product sound is primary, describe the physical action that creates it. If music is primary, state the tempo, instrumentation, mood, and where the arrangement should change.

3. Add ambience as an acoustic place

Ambience should explain where the shot lives. Instead of “café sounds,” describe a quiet espresso-machine hiss behind the speaker, low room conversation, and a door chime in the distance. For a clean product spot, ask for controlled studio room tone rather than silence unless silence is a deliberate creative choice.

Use distance and material words that can apply to both picture and sound:

  • close, dry, intimate, and lightly reverberant
  • distant, muffled through glass, or off-camera left
  • footsteps on tile, fabric movement, rain on metal, or a glass set on wood

4. Decide whether music is score or source

Score sits outside the scene; source music comes from an object or location inside it. State which one you mean. “A restrained electronic score under the dialogue” asks for a background layer, while “music playing quietly from the café radio, slightly muffled” gives it an on-screen acoustic source.

When music is not needed, say no music. This keeps a dialogue or Foley test easier to judge. When it is needed, leave headroom for the voice instead of asking every layer to be loud and dramatic.

5. State the exclusions

Negative audio direction is useful when it protects the brief. Examples include no narration, no crowd cheering, no added melody, and no exaggerated whooshes. Keep the exclusion list short; too many constraints can compete with the action you actually want.

How to use audio references without confusing their role

H3's reference system is most useful when every asset has one job. A reference audio clip might supply vocal timbre and cadence, musical energy, a sound texture, or editing rhythm. Do not assume the model knows which property to borrow.

Write an asset map before the final prompt:

AssetAssigned jobWhat should not transfer
Character imageIdentity, wardrobe, and product placementBackground and lighting
Motion videoGesture and camera timingCharacter identity and dialogue
Voice audioVocal character, cadence, and emotional energyOriginal words and background noise
Music audioTempo, arrangement density, and edit rhythmRecognizable melody or lyrics

Then describe those relationships in natural language. MiniMax's own H3 example uses a sentence that assigns camera motion to a video reference, the performer to an image reference, and vocals to an audio reference. Numbered asset labels vary by interface, while the API uses typed content and reference roles, so replace the labels in the templates below with the exact notation shown in your workflow.

Warning

Only use a person's voice or performance as a reference when you have permission. A voice reference is not evidence that you own the speaker's identity, the recording, the music, or the right to publish a generated imitation.

Clean references make cleaner experiments. Trim silence, remove unrelated background music, avoid clipping, and keep the relevant vocal or musical passage easy to hear. Use a short segment that demonstrates the property you want rather than uploading the maximum duration by default.

Five MiniMax H3 native audio prompt templates

These are starting briefs, not claimed test results. Run each prompt more than once and score the outputs using the protocol in the next section.

1. Single-speaker dialogue with room tone

Close-up of a chef at the pass in a quiet restaurant kitchen after service. The camera makes a slow push-in as she looks toward a colleague off-camera and says, “We try it again tomorrow.” Delivery: tired but quietly optimistic, natural pace, one short breath before “tomorrow.” Her voice is the primary audio layer, close and dry. Low ventilation hum and occasional metal cooling sounds remain distant. No music, no narration, no other voices.

Use this template to evaluate speech accuracy, emotional delivery, mouth movement, and whether the ambience stays behind the line.

2. Layered ambience with synchronized Foley

Wide shot inside a nearly empty laundromat at night, fluorescent light reflecting on the tiled floor. A person moves a wet jacket from a washer to a rolling basket; the metal door clicks shut exactly when their hand closes it, then the basket wheels rattle softly across tile. Foreground: fabric movement and the door click. Mid-ground: steady washer rotation. Background: rain against the front window and one distant car passing outside. Native stereo mix, no dialogue, no music.

This brief isolates ambience, material realism, left-right separation, and contact synchronization.

3. Music-led product reveal

Ten-second vertical product film for a translucent coral sports bottle on a lilac studio surface. Start with a macro shot of condensation, then orbit as the bottle rises into the hero position. Original minimal electronic score at 100 BPM: soft pulse for the opening, a clean percussive accent when the logo faces camera, then a short resolved final note. Add subtle droplets and one controlled placement sound. Music supports the product sounds rather than masking them. No voice and no stock-style cinematic boom.

For product-specific visual planning beyond the soundtrack, use the photo-to-product-video workflow.

4. Permitted voice reference with new dialogue

Use voice reference 1 only for the narrator's vocal timbre, speaking pace, and warm conversational energy. Do not reuse its words or background noise. The person from image reference 1 stands beside the window and says, “The first draft is only the beginning.” Keep the spoken line intelligible and synchronized to the visible speaker. Add a quiet residential room tone and soft rain outside. No music. The voice recording is authorized for this use.

Test whether the intended vocal qualities carry over without unintentionally copying noise, phrasing, or content from the source recording. Describe the desired qualities in words so the instruction remains useful even if reference adherence varies.

5. Rhythm reference for a social edit

Create a 12-second montage of a ceramic artist shaping, glazing, and placing a finished cup on a shelf. Use audio reference 1 only for tempo, rhythmic energy, and edit pacing. Compose an original percussive track with different melody and sound design. Cut the visual action on the major rhythmic changes: clay lands on the wheel, brush touches glaze, cup reaches the shelf. Preserve natural wheel rotation, brush texture, and the final ceramic tap beneath the music. No dialogue.

This separates rhythmic conditioning from music copying and gives you three visible sync points to inspect.

A repeatable H3 audio evaluation plan

One polished official demo cannot answer whether a model fits your production workflow. Use the same brief for at least three generations, keep duration and references fixed, and score every output rather than selecting only the strongest result.

Test the layers separately first

Begin with four controlled prompts:

  1. Dialogue and quiet room tone only
  2. Ambience and three visible Foley events, with no speech or music
  3. Instrumental music with two planned visual accents
  4. One authorized audio reference paired with a simple image or video reference

Only after those pass should you combine dialogue, ambience, Foley, and music in one crowded scene. This makes failures diagnosable: you can tell whether a problem comes from pronunciation, timing, reference transfer, or mix density.

Use a scorecard, not an impression

DimensionQuestion to answerSimple measure
Dialogue accuracyWere the requested words spoken correctly?Correct words divided by requested words
Lip-syncDo visible mouth closures and syllable stresses align?Score 1–5 and note the first visible drift
Event syncDo contact sounds land with the pictured action?Frames early or late at three planned cues
Voice consistencyDoes the speaker remain recognizable through the line?Score beginning, middle, and end separately
Reference controlDid the requested property transfer without unwanted baggage?List intended and unintended transfers
Stereo imageAre left-right placements useful and stable?Headphone check plus channel waveform check
Mix hierarchyIs the primary layer clearly intelligible?Score 1–5 at normal playback level
RepeatabilityHow often is the clip usable without audio repair?Usable outputs divided by total runs

Inspect the exported file in a video editor or audio workstation. Confirm that two channels exist, listen for phase problems in mono, and compare waveforms around planned contact events. Do not infer audio quality from a social platform preview, which may recompress or remap the soundtrack.

Tip

Keep the failed generations. A prompt revision is only an improvement if it raises the usable-output rate across repeated runs, not if it produces one lucky clip.

Limitations to plan around

MiniMax's own launch post says H3 still has room to improve and identifies visual detail as one area for future work. For audio production, the current public documentation also leaves several questions open. It does not publish guarantees for word error rate, lip-sync accuracy, loudness, sample rate, channel layout beyond stereo, language coverage, voice-reference similarity, or timecode-level cue control.

Plan around these practical boundaries:

  • Short output: 4–15 seconds is suitable for shots, ads, and social clips, but long-form continuity needs editing across generations.
  • Short reference context: Up to three audio clips are allowed, but their combined duration cannot exceed 15 seconds.
  • No audio-only reference request: A reference audio clip must be paired with at least one image or video.
  • Natural-language timing is not a locked timeline: Treat beat timings as intent until repeated tests prove the precision you need.
  • Stereo needs verification: Two channels do not automatically create believable direction, depth, or mono compatibility.
  • Dense mixes can hide errors: Dialogue, effects, ambience, and music may all be modeled together, but a simpler brief is easier to control and repair.
  • Reference rights still apply: Model capability does not grant permission to imitate a voice or reuse protected music.

Availability, interface controls, and billing can also change. MiniMax's current official pricing lists audio reference material without a separate input charge, while video output and some other reference materials are billed under their own rules. Check the official pay-as-you-go pricing page before budgeting an API production; do not substitute a third-party host's rate for MiniMax's own price. The H3 pricing and specs guide works through MiniMax's official rates and reference-material costs. The current OmniArt model catalog does not establish H3 availability, so this guide does not claim that H3 can already be generated inside OmniArt.

Build the rest of the soundtrack in OmniArt

Native audio can reduce handoffs, but it does not remove the value of separate sound tools. If an H3 take has the right picture and weak narration, keep the picture and build a controlled voice track. If the dialogue works but the room lacks texture, add an ambience or Foley layer instead of regenerating the complete shot.

OmniArt brings video, speech, sound, and music workflows into one creative workspace. Use the AI sound effect generator guide to design replacement Foley and ambience, or compare the prompting approach with the Veo 3.1 spatial audio guide. The goal is not to force every sound into one generation. It is to keep the native track when it works and replace only the layer that does not.

Explore video creation workflows on OmniArt

Ready to Create?

Start generating amazing content with AI

Get started free