guide튜토리얼 및 사용 가이드26분 읽기

MiniMax H3 프롬프트 가이드: 모든 참조 입력 제어하기

이미지, 영상, 오디오, 텍스트 참조를 위한 MiniMax H3 프롬프트 작성법을 알아보세요. 에셋별 역할을 정하고 충돌을 막아 시간 지정 AI 영상 샷을 연출합니다.

OmniArt 팀
MiniMax H3 프롬프트 가이드: 모든 참조 입력 제어하기

MiniMax H3 프롬프트의 핵심은 “더 자세히 쓰기”가 아니라 “모든 입력에 하나의 역할 주기”입니다. H3는 텍스트, 이미지, 영상, 오디오를 통합 참조 세트로 받을 수 있으므로, 크리에이터는 정지 이미지에서 제품 정체성을, 클립에서 움직임을, 오디오에서 페이싱을, 프롬프트에서 샷 연출을 가져올 수 있습니다. 이 유연성은 모델이 어떤 출처가 결과의 어느 부분을 제어하는지 알 때에만 도움이 됩니다.

이 MiniMax H3 프롬프트 가이드는 그 원칙을 반복 가능한 형식으로 바꿉니다. 각 에셋을 역할에 연결하고, 참조 간 모순을 막고, 보존 규칙을 작성하고, 4–15초 생성을 시간별 비트로 나누는 방법을 배웁니다. 마지막의 6개 템플릿은 제품 영상, 캐릭터 모션, 대사, 영상 편집, UI 모션, 첫 프레임에서 마지막 프레임으로의 전환을 다룹니다.

참고

이 글은 OmniArt가 테스트한 H3 생성 결과 보고서가 아니라 문서 기반 프롬프트 가이드입니다. MiniMax의 최신 H3 생성 가이드공식 API 참조를 바탕으로 합니다. 2026년 7월 31일 기준 H3는 OmniArt의 사실 기준 모델 카탈로그에 아직 포함된 것으로 확인되지 않았습니다. 인터페이스 레이블과 참조 문법도 MiniMax 자체 서비스와 하위 호스트에서 다를 수 있습니다.

MiniMax H3, Hailuo 3.0, Hailuo 03 중 무엇인가요?

MiniMax의 현재 문서는 이 모델을 MiniMax H3라고 부르며 공식 API 모델 식별자는 MiniMax-H3입니다. Hailuo는 MiniMax의 영상 제품 브랜드이므로 출시 보도와 제3자 카탈로그에서는 Hailuo 3.0, Hailuo 03, Hailuo-3.0도 볼 수 있습니다.

현재 사용하는 서비스가 요구하는 레이블을 쓰세요. 게시 시점의 공식 MiniMax V2 API에서는 MiniMax-H3를 뜻하며, 마켓플레이스 별칭이 공식 API 식별자로도 유효하다고 가정해서는 안 됩니다. 이 가이드는 모델에는 MiniMax H3, 이미지·영상·오디오 입력을 결합하는 모드에는 참조 생성을 사용합니다. “옴니 참조”는 이 혼합 미디어 워크플로를 설명하는 데 유용하지만, 공식 API 명칭은 reference-to-video입니다.

세 가지 H3 생성 모드 이해하기

프롬프트를 쓰기 전에 올바른 모드를 선택해야 합니다.

모드입력적합한 경우
텍스트-영상필수 텍스트 프롬프트설명만으로 장면을 만들 수 있을 때
첫/마지막 프레임 이미지-영상필수 프롬프트와 첫 프레임, 마지막 프레임 또는 둘 다정확한 시작 또는 종료 구성이 중요할 때
참조 생성필수 프롬프트와 참조 이미지, 영상 또는 오디오업로드 에셋의 정체성, 디자인, 움직임, 카메라, 음성, 리듬 또는 스타일이 필요할 때

현재 공식 사양은 2K 출력과 4–15초의 정수 길이를 제시합니다. 모든 요청에는 비어 있지 않은 텍스트 프롬프트가 필요하고 제한은 7,000자입니다. 참조 요청은 이미지 최대 9개, 영상 3개, 오디오 클립 3개를 포함할 수 있으며 전체 에셋 수는 12개를 넘을 수 없습니다. 참조 영상과 참조 오디오도 각각 합계 15초로 제한됩니다. 오디오 참조는 단독 제출할 수 없고 최소 하나의 이미지나 영상을 함께 사용해야 합니다.

텍스트-영상에는 구체적인 화면비가 필요합니다. 참조 생성은 구체적인 비율을 사용하거나 H3가 적응적으로 선택하게 둘 수 있고, 첫/마지막 프레임 생성은 업로드한 이미지를 따릅니다. 문서화된 비율은 21:9, 16:9, 4:3, 1:1, 3:4, 9:16입니다.

확실한 워크플로 경계가 하나 있습니다. 첫/마지막 프레임 모드와 참조 생성은 같은 API 요청에서 섞을 수 없습니다. 정확한 경계 프레임이 중요하면 첫/마지막 프레임 모드를, 정체성·움직임·스타일·오디오 전송이 더 중요하면 참조 생성을 사용하고 시작 구성을 텍스트로 설명하세요.

역할 맵으로 H3 프롬프트 만들기

H3 API는 reference_image, reference_video, reference_audio 같은 역할이 지정된 항목으로 미디어를 받습니다. 이 기술 역할은 미디어 유형을 식별할 뿐이므로, 프롬프트에서는 각 에셋의 창의적 책임을 한 단계 더 명시해야 합니다.

다음 순서를 사용하세요.

DELIVERABLE
What the final clip is for, its duration, and its format.

ROLE MAP
Reference image 1 = subject identity and immutable appearance.
Reference image 2 = environment or visual style only.
Reference video 1 = body motion and timing only.
Reference audio 1 = rhythm and edit timing only.

SHOT
Subject, action, environment, lighting, and camera direction.

TIMELINE
Time-coded beats that add up to the selected duration.

PRESERVE
Features that must remain unchanged from named references.

EXCLUDE
A short list of the most damaging unwanted changes.

위 번호 레이블은 브리프를 위한 의미상의 레이블이지, 모든 H3 인터페이스가 문자 그대로 Reference image 1 핸들을 제공한다는 뜻은 아닙니다. 공식 API에서는 파일이 content 배열의 별도 항목입니다. 시각적 인터페이스에서는 해당 서비스가 제공하는 첨부 이름이나 언급 문법을 사용하되, 하나의 에셋에 하나의 역할을 주는 논리는 유지하세요.

각 에셋에 하나의 주요 역할 주기

참조에는 보통 전송하고 싶지 않은 정보도 더 많이 담겨 있습니다. 댄스 클립에는 연기자, 의상, 장소, 카메라, 조명, 움직임, 음악까지 있을 수 있습니다. H3에 “영상을 따르라”고만 하면 그 모든 특성이 캐릭터 이미지와 경쟁합니다.

대신 좁게 연결하세요.

  • 이미지 참조: 정체성, 얼굴, 의상, 제품 형상, 로고, 팔레트 또는 환경 디자인
  • 영상 참조: 신체 움직임, 카메라 경로, 동작 타이밍, 물리적 상호작용 또는 편집 리듬
  • 오디오 참조: 음성 특성, 억양, 음악 리듬, 분위기음 또는 이벤트 큐
  • 텍스트 프롬프트: 최종 장면, 전송 규칙, 타임라인, 우선순위 및 보존 지시

한 에셋이 두 역할을 해야 한다면 둘 다 이름을 붙이고 경계를 명확히 하세요. 예를 들어 “참조 영상 1은 카메라 경로와 동작 타이밍을 제어한다. 연기자, 의상, 장소, 색보정, 오디오는 무시한다”라고 씁니다.

참조끼리 충돌하지 않게 하기

참조 충돌은 대개 모델 문제보다 먼저 지시 문제입니다. 두 이미지에 서로 다른 재킷이 있거나, 모션 클립의 체형이 다르거나, 오디오 비트가 요청한 끊김 없는 카메라 이동과 충돌할 수 있습니다.

입력이 겹칠 때는 서면 우선순위를 사용하세요.

Priority order:
1. Reference image 1 controls character identity and wardrobe.
2. Reference video 1 controls movement and timing only.
3. Reference image 2 controls lighting palette and set design only.
4. The text prompt controls camera framing and final composition.

그다음 H3가 선호를 추론하길 바라지 말고 모순을 해결하세요. 정체성 이미지에는 긴 머리가 있고 모션 영상에는 모자가 있다면, 최종 캐릭터가 머리를 유지할지 모자를 쓸지 명시하세요. 제품 참조에 읽을 수 있는 라벨이 있지만 스타일 이미지가 회화적이라면, 스타일화는 환경에만 적용하고 패키지는 사실적으로 유지한다고 말하세요.

세 가지 습관이 충돌을 줄입니다.

  1. 가장 작은 유용한 참조 세트를 사용하세요. 입력이 많을수록 불일치 가능성도 커집니다. 프롬프트가 안정적으로 표현할 수 없는 정보를 제공할 때만 파일을 추가하세요.
  2. 콘텐츠와 스타일을 분리하세요. 어느 참조가 피사체를, 어느 참조가 표현 방식을 소유하는지 이름을 붙이세요.
  3. 카메라 권한을 하나 선택하세요. 영상에서 카메라를 전송하거나 텍스트로 연출하세요. 둘 다 필요하다면 영상은 타이밍을 제공하고 텍스트가 프레이밍을 덮어쓴다고 명시하세요.

더 폭넓은 증상별 워크플로는 반복 작업 중 AI 영상 프롬프트 수정 7가지를 이 가이드와 함께 보세요.

확인 가능한 보존 지시 작성하기

“일관되게 유지”는 너무 모호합니다. 보존 블록은 어떤 프레임에서도 검사할 수 있는 가시적 속성을 이름으로 지정해야 합니다.

캐릭터의 경우:

Preserve from reference image 1 throughout every frame: facial structure,
eye color, hairstyle and length, jacket cut, jacket color, and body proportions.
Do not replace the performer with the person from reference video 1.

제품의 경우:

Preserve from reference image 1: bottle silhouette, cap shape, label placement,
navy wordmark, coral seal, material finish, and relative proportions.
No duplicate product, redesigned packaging, extra text, or changing logo.

긍정적 앵커를 먼저 두고 제외 항목은 난간으로 사용하세요. 부정 목록 전체가 되어서는 안 됩니다. 모호한 금지 10개는 결과 사용 가능성을 실제로 결정하는 불변 요소 4개를 희석할 수 있습니다.

같은 캐릭터가 별도로 생성한 여러 클립에 등장한다면, 모든 프롬프트에서 정체성과 의상 블록을 동일하게 유지하세요. 일관된 캐릭터 워크플로는 더 긴 시퀀스 전체에 재사용할 참조 패킷을 만드는 방법을 설명합니다.

콘텐츠만이 아니라 시간도 지시하기

H3는 4–15초 클립을 허용하지만 긴 문단만으로는 각 이벤트가 언제 발생하는지 모델에 알려 주지 못합니다. 타임라인은 프롬프트를 간결한 샷 계획으로 바꿉니다.

10초 클립의 예시는 다음과 같습니다.

0–2 seconds: locked medium-wide establishing shot; subject holds still.
2–6 seconds: subject performs the referenced action at the source tempo.
6–9 seconds: camera makes one slow 20-degree arc to the right.
9–10 seconds: subject settles; hold a clean final composition for the edit.

각 비트는 달성 가능하게 유지하세요. 주요 동작 하나와 카메라 움직임 하나가 관련 없는 변형을 줄줄이 잇는 것보다 대개 명확합니다. 시간 범위의 합이 선택한 길이가 되게 하고, 중요한 제품 또는 얼굴 디테일은 느린 비트에 두며, 마지막 0.5초 또는 1초는 안정적인 편집 지점으로 남기세요.

오디오도 같은 시계를 공유할 수 있습니다.

At 2.0 seconds, the first downbeat starts the hand movement.
At 6.0 seconds, the bass hit motivates the camera arc.
From 9.0 seconds, let the music tail continue under the held final frame.

렌즈, 조명, 카메라 어휘는 시네마틱 AI 영상 프롬프트 가이드를 참조하세요.

여섯 가지 MiniMax H3 프롬프트 템플릿

대괄호 안 세부 정보를 바꾸고, 열거된 에셋만 첨부하며, 레이블은 사용하는 인터페이스에 맞추세요. 이 템플릿은 문서화된 H3 입력 모드를 기반으로 한 구조화된 출발점이며, 벤치마크 우승 결과나 보장된 출력이라는 주장은 아닙니다.

템플릿 1: 모션 및 음악 참조가 있는 제품 공개

  • 모드: 참조 생성
  • 에셋: 깨끗한 제품 이미지 1개, 카메라 모션 영상 1개, 선택적 음악 참조
Create a 10-second 9:16 product reveal for a paid social placement.

Role map:
- Reference image 1 controls the product's exact design, proportions, packaging,
  colors, materials, label placement, and wordmark.
- Reference video 1 controls camera path and acceleration only. Ignore its subject,
  location, lighting, color grade, and audio.
- Reference audio 1 controls beat and edit timing only.

Scene: The product stands centered on a warm-violet studio plinth. Soft lilac key
light from camera left, restrained coral rim light, subtle atmospheric haze.

Timeline:
- 0–2 seconds: static wide reveal on the first soft beat.
- 2–7 seconds: transfer the smooth push-and-arc movement from reference video 1.
- 7–9 seconds: one narrow highlight travels across the product surface.
- 9–10 seconds: camera settles; hold the label front-facing and readable.

Preserve reference image 1 exactly: silhouette, cap, material finish, label layout,
brand colors, and wordmark. Keep one product only. No redesigned packaging,
duplicate objects, extra text, warped label, or abrupt camera shake.

이 프롬프트 전에 필요한 정지 이미지 준비는 사진을 제품 영상으로 만드는 워크플로를 사용하세요.

템플릿 2: 정체성 드리프트 없는 캐릭터 모션 전송

  • 모드: 참조 생성
  • 에셋: 캐릭터 이미지 1개, 모션 영상 1개, 선택적 환경 이미지
Create an 8-second 16:9 cinematic character performance.

Priority order:
1. Reference image 1 controls face, hair, body proportions, and wardrobe.
2. Reference video 1 controls body choreography and action timing only.
3. Reference image 2 controls the environment palette and architecture only.
4. This prompt controls framing and lighting.

The character from reference image 1 performs the complete movement from reference
video 1 in a moonlit station based on reference image 2. Medium full shot, camera
locked at chest height, soft directional light, realistic weight and foot contact.

Timeline:
- 0–1 seconds: character holds the starting pose.
- 1–7 seconds: perform the reference choreography once at its original tempo.
- 7–8 seconds: settle naturally and look toward camera left.

Preserve facial structure, hairstyle, coat shape, coat color, boots, and body
proportions from reference image 1 in every frame. Do not inherit the motion video's
performer, face, clothing, background, camera movement, or audio. No extra limbs,
sliding feet, costume changes, or cuts.

템플릿 3: 음성 참조를 사용한 대사 퍼포먼스

  • 모드: 참조 생성
  • 에셋: 캐릭터 이미지 1개, 허가받은 음성 또는 대사 오디오 클립 1개

경고

소유하거나 사용할 명시적 허가를 받은 음성만 업로드하세요. 모델이 참조 오디오를 받아들인다고 해서 다른 사람의 음성을 모방할 권리가 생기지는 않습니다.

Create a 9-second 16:9 single-character dialogue shot.

Role map:
- Reference image 1 controls the speaker's appearance, wardrobe, and room design.
- Reference audio 1 controls spoken timing, cadence, emotional progression, and pauses.

Shot: Medium close-up, eye-level, 50 mm cinematic framing. The speaker begins calm,
briefly smiles after the central pause, then finishes with quiet confidence. Natural
blinks and restrained hand movement. Camera remains static; soft room tone underneath.

Timeline:
- 0–1 seconds: silent eye contact and a small inhale.
- 1–8 seconds: performance follows reference audio 1 exactly in timing and pauses.
- 8–9 seconds: mouth closes, expression settles, hold for the edit.

Preserve face, hairstyle, skin tone, jacket, background layout, and lighting direction
from reference image 1. Keep one speaker. No camera move, cutaway, background speech,
new words, exaggerated gestures, or wardrobe changes.

템플릿 4: 원본 영상의 환경 교체

  • 모드: 참조 생성
  • 에셋: 원본 영상 1개, 환경 이미지 1개, 선택적 피사체 이미지
Create a 12-second 16:9 environmental replacement based on the source clip.

Role map:
- Reference video 1 controls shot length, subject action, physical timing, camera path,
  framing progression, and interaction with the ground.
- Reference image 1 controls the replacement environment, architecture, palette,
  weather, and lighting mood.
- Reference image 2 controls the subject's face and wardrobe, if supplied.

Replace the source video's location with the environment from reference image 1.
Preserve the original action and camera movement continuously. Match subject lighting,
contact shadows, reflections, and atmospheric perspective to the new environment.

Timeline follows reference video 1. Maintain one continuous shot with the original
action beats and no added event.

Preserve the source subject's position, scale, motion, and ground contact. Preserve
identity and wardrobe from reference image 2 when present. Do not copy the original
background, signage, bystanders, color grade, or source audio. No cuts, teleporting,
floating feet, or changing architecture.

템플릿 5: 비트에 동기화된 브랜디드 UI 모션

  • 모드: 참조 생성
  • 에셋: UI 레이아웃 이미지 1개, 모션 스타일 영상 1개, 선택적 오디오 참조
Create a 7-second 1:1 UI motion-design clip for a product announcement.

Priority order:
1. Reference image 1 controls exact layout, hierarchy, component positions, text,
   logo, colors, corner shapes, and typography appearance.
2. Reference video 1 controls transition character and easing only.
3. Reference audio 1 controls the timing of three motion beats only.

Animate the interface from reference image 1 without redesigning it. Begin with the
main panel at rest. Use the restrained slide-and-scale transition quality from
reference video 1. Keep the camera orthographic and the background static.

Timeline:
- 0–1 seconds: complete layout at rest.
- 1–3 seconds: cards enter in sequence on beat one.
- 3–5 seconds: primary control changes state on beat two.
- 5–6 seconds: one subtle emphasis pulse on beat three.
- 6–7 seconds: complete interface holds sharp and readable.

Preserve every word, logo shape, component proportion, spacing relationship, and brand
color from reference image 1. No invented labels, misspelled text, extra panels,
perspective tilt, camera movement, glow overload, or elastic distortion.

템플릿 6: 정확한 첫 프레임에서 마지막 프레임 변환

  • 모드: 첫/마지막 프레임 이미지-영상
  • 에셋: 첫 프레임 이미지 1개와 마지막 프레임 이미지 1개, 참조 미디어 없음
Create a 10-second transformation from the supplied first frame to the supplied last
frame. Preserve the subject's identity and the camera's fixed position throughout.

Timeline:
- 0–2 seconds: hold the first-frame composition; only subtle ambient movement.
- 2–7 seconds: the scene transforms progressively from the center outward. Materials
  change continuously with believable physical contact and no hard cut.
- 7–9 seconds: remaining details resolve into the last-frame design.
- 9–10 seconds: arrive exactly at the supplied last frame and hold it cleanly.

One continuous locked shot. Preserve subject scale, face, silhouette, horizon, lens,
and framing during the transition. No reference-style transfer, new characters,
camera movement, jump cut, flicker, or overshoot beyond the final composition.

이 요청에는 다른 템플릿의 참조 이미지, 영상, 오디오를 첨부하지 마세요. 공식 API는 첫/마지막 프레임 생성과 참조 생성을 상호 배타적인 모드로 취급합니다.

실용적인 반복 순서

한 번의 만족스럽지 않은 렌더 뒤에 정체성, 움직임, 카메라, 스타일, 오디오 지시를 모두 바꾸지 마세요. 그러면 어떤 지시가 도움이 되었는지 알 수 없습니다.

세 단계로 반복하세요.

  1. 역할과 불변 요소를 고정합니다. 가장 작은 참조 세트를 사용하고 피사체나 제품이 알아볼 수 있게 유지되는지 확인합니다.
  2. 동작과 타임라인을 고정합니다. 스타일은 단순하게 유지한 채 모션 출처 또는 시간 비트를 추가합니다.
  3. 표현을 더합니다. 앞의 두 층이 안정된 뒤 환경, 조명, 색보정, 오디오 큐, 마무리 디테일을 도입합니다.

결과가 드리프트하면 더 많은 부정적 언어를 추가하기 전에 단순화하세요. 중요도가 가장 낮은 참조를 제거하고, 타임라인을 줄이거나, 한 출처에 더 좁은 역할을 부여하세요. 정체성, 형상, 움직임, 카메라, 오디오 타이밍, 마지막 프레임 안정성이라는 같은 점검표로 여러 변형을 비교하세요.

OmniArt에서 시작하기

MiniMax H3는 현재 OmniArt에서 선택 가능한 모델로 확인되지 않았으므로, 이 글은 H3 생성 경로를 연결하거나 OmniArt H3 워크플로를 주장하지 않습니다. 현재 제공되는 모델은 현재 OmniArt 영상 모델 라인업에서 확인하세요.

역할 맵 방법은 지원되는 참조 기반 영상 모델에도 잘 적용됩니다. 작업 공간을 열기 전에 정체성 이미지 하나, 모션 출처 하나, 짧은 보존 블록을 준비하세요. 창작 브리프가 안정될 때까지 프롬프트는 모델 중립적으로 유지한 다음, 실제 선택한 모델에 맞춰 첨부 레이블과 기능별 문법을 조정하세요. 그러면 지금은 재사용 가능한 연출 시트를 얻고, 나중에 모델이 OmniArt에 추가되면 깔끔한 H3 프롬프트도 얻게 됩니다.

공식 API 요금, 참조 영상 과금 및 계산 예시는 MiniMax H3 가격 및 사양 가이드를 사용하세요. 리셀러 요금이 아니라 MiniMax 자체 가격 페이지를 기준으로 합니다.

제작할 준비가 되셨나요?

AI로 멋진 콘텐츠를 생성하세요

무료로 시작하기