MiniMax H3 提示詞指南:管好每一份參考輸入
學會 MiniMax H3 在圖片、影片、音訊和文字參考上的提示詞寫法。給每份素材一個角色、避免互相打架,並排出有時間軸的鏡頭。

MiniMax H3 提示詞背後真正有用的想法,不是「多寫一點細節」,而是「給每一份輸入一個任務」。H3 可以把文字、圖片、影片和音訊當成一組統一的參考素材,所以創作者能從一張靜態圖拿產品身分、從一段片子拿動作、從音訊拿節奏、從提示詞拿鏡頭調度。但這種彈性只有在模型分得出哪個來源負責結果的哪一部分時才幫得上忙。
這份 MiniMax H3 提示詞指南把那個原則變成一套可以重複使用的格式。你會學到怎麼把每份素材綁到一個角色上、怎麼避免參考素材彼此矛盾、怎麼寫保留規則,以及怎麼把一段 4–15 秒的生成拆成有時間點的節拍。文末六組範本涵蓋產品影片、角色動作、對白、影片改景、UI 動態,以及首尾影格轉場。
說明
這是一份以官方文件為依據的提示詞指南,不是 OmniArt 獨立實測 H3 生成 結果的報告。它依據的是 MiniMax 目前的 H3 生成指南 和官方 API 說明文件。 MiniMax H3 現在已經開放給 OmniArt 的 Creator 以上使用者,但服務商文件裡 記載的能力,仍然不該被讀成獨立的實測結果。
MiniMax H3、Hailuo 3.0 還是 Hailuo 03?
MiniMax 目前的文件把這個模型稱為 MiniMax H3,官方 API 的模型識別碼是 MiniMax-H3。你在發布報導和第三方目錄裡,可能也會看到 Hailuo 3.0、Hailuo 03 或 Hailuo-3.0,因為 Hailuo 是 MiniMax 的影片產品品牌。
請照你眼前那個介面要求的名稱來用。以本文發布時的官方 MiniMax V2 API 來說,那就是 MiniMax-H3;不要假設某個平台的別名同時也是合法的官方 API 識別碼。本文用 MiniMax H3 指這個模型,用參考生成指那個把圖片、影片和音訊輸入混在一起的模式。「Omni reference」是形容這種混合素材流程的好說法,但官方 API 把它叫做 reference-to-video。
先搞懂 H3 的三種生成模式
挑對模式,比寫提示詞更早。
| 模式 | 輸入 | 什麼時候用 |
|---|---|---|
| 文字生成影片 | 必填的文字提示詞 | 場景可以完全從描述長出來 |
| 首尾影格圖生影片 | 必填提示詞,加上首影格、尾影格或兩者 | 開頭或結尾的構圖必須精準 |
| 參考生成 | 必填提示詞,加上參考圖片、影片或音訊 | 你需要從上傳素材取得身分、設計、動作、運鏡、聲音、節奏或風格 |
目前的官方規格寫的是 2K 輸出,以及 4 到 15 秒的整數時長。每一次請求都需要非空白的文字提示詞,上限 7,000 個字元。一次參考請求最多可以包含 9 張圖片、3 段影片和 3 段音訊,素材總數不超過 12 份。參考影片總長上限 15 秒,參考音訊也一樣。音訊參考不能單獨送出,必須至少搭配一張圖片或一段影片。
文字生成影片必須指定明確的畫面比例。參考生成可以指定明確比例,也可以讓 H3 自行決定;首尾影格生成則跟著上傳的圖片走。文件記載的比例選項是 21:9、16:9、4:3、1:1、3:4 和 9:16。
有一條硬性的流程界線:首尾影格模式和參考生成不能放在同一次 API 請求裡。如果精準的邊界影格很重要,就用首尾影格模式;如果身分、動作、風格或聲音的移轉更重要,就用參考生成,並用文字描述開場構圖。
把 H3 提示詞寫成一張角色分工表
H3 的 API 接收的媒體是有型別的項目,角色像是 reference_image、reference_video 和 reference_audio。那些技術角色標示的是媒體型別;你的提示詞要再往前一步,講明每份素材在創作上負責什麼。
用這個順序:
DELIVERABLE
What the final clip is for, its duration, and its format.
ROLE MAP
Reference image 1 = subject identity and immutable appearance.
Reference image 2 = environment or visual style only.
Reference video 1 = body motion and timing only.
Reference audio 1 = rhythm and edit timing only.
SHOT
Subject, action, environment, lighting, and camera direction.
TIMELINE
Time-coded beats that add up to the selected duration.
PRESERVE
Features that must remain unchanged from named references.
EXCLUDE
A short list of the most damaging unwanted changes.
上面那些編號標籤,是給你自己的需求文件用的語意標籤,不代表每個 H3 介面都真的開放 Reference image 1 這種代號。在官方 API 裡,這些檔案是 content 陣列裡各自獨立的項目。在圖形介面裡,就用那個介面提供的附件名稱或提及語法,然後維持同一套「一份素材一個任務」的邏輯。
每份素材只給一個主要任務
一份參考素材裡的資訊,通常比你想移轉的多。一段舞蹈影片裡有表演者、衣服、場地、鏡頭、光線、動作,可能還有音樂。如果你叫 H3「跟著那段影片走」,這些特徵全部都會和你的角色圖搶位置。
要綁就綁得窄:
- 圖片參考: 身分、臉、服裝、產品結構、標誌、色彩,或環境設計。
- 影片參考: 身體動作、運鏡路徑、動作時間點、物理互動,或剪接節奏。
- 音訊參考: 聲音特質、語調節奏、音樂節拍、環境音,或某個事件提示。
- 文字提示詞: 最終場景、移轉規則、時間軸、優先順序和保留指示。
當某份素材非得做兩件事時,兩件都講明,並把界線寫清楚:「Reference video 1 controls camera path and action timing; ignore its performer, wardrobe, location, color grade, and audio.」
別讓參考素材互相打架
參考素材衝突,通常先是指示的問題,才是模型的問題。兩張圖片可能穿著不同的外套、動作片段可能是不同體型的人,或者音訊的節拍暗示了剪接點,卻和你要求的不中斷運鏡衝突。
只要輸入有重疊,就寫下優先順序:
Priority order:
1. Reference image 1 controls character identity and wardrobe.
2. Reference video 1 controls movement and timing only.
3. Reference image 2 controls lighting palette and set design only.
4. The text prompt controls camera framing and final composition.
接著把矛盾解決掉,不要指望 H3 會自己猜出你的偏好。如果身分圖裡是散著的頭髮、動作影片裡卻戴著帽子,就說清楚最終角色是留頭髮還是戴帽子。如果產品參考圖上的標籤清晰可讀、風格圖卻是繪畫感的,就說明風格化只套用在環境上,包裝維持寫實。
有三個習慣能減少衝突:
- 用最小的有效素材組。 輸入越多,可能的分歧就越多。只有當某個檔案能提供提示詞無法穩定表達的資訊時,才把它加進去。
- 把內容和風格分開。 講明哪份參考擁有主體,哪份擁有處理方式。
- 只留一個運鏡的話事者。 要嘛從影片移轉運鏡,要嘛用文字指揮。如果兩者都要,就說明影片負責時間點、文字覆寫取景。
想要一份按症狀對症下藥的流程,迭代時可以把 7 個 AI 影片提示詞修正法放在這份指南旁邊。
寫出可以檢查的保留指示
「保持一致」太含糊了。一段保留規則應該點名那些你能在任何一格畫面上檢查的可見屬性。
角色的寫法:
Preserve from reference image 1 throughout every frame: facial structure,
eye color, hairstyle and length, jacket cut, jacket color, and body proportions.
Do not replace the performer with the person from reference video 1.
產品的寫法:
Preserve from reference image 1: bottle silhouette, cap shape, label placement,
navy wordmark, coral seal, material finish, and relative proportions.
No duplicate product, redesigned packaging, extra text, or changing logo.
正面的錨點要放在前面;排除項是護欄,不是提示詞的全部。負面清單要短而具體。十條含糊的禁令,會稀釋掉真正決定這一條能不能用的那四個不變項。
如果同一個角色要出現在好幾段分開生成的片子裡,就讓身分和服裝那一段在每條提示詞裡都一模一樣。角色一致性流程說明了怎麼在較長的序列裡建立這種可重複使用的參考素材包。
指揮時間,不只是內容
H3 允許 4 到 15 秒的片段,但一大段文字不會告訴模型每個事件發生在什麼時候。時間軸把提示詞變成一份精簡的鏡頭計畫。
以一段 10 秒的片子為例:
0–2 seconds: locked medium-wide establishing shot; subject holds still.
2–6 seconds: subject performs the referenced action at the source tempo.
6–9 seconds: camera makes one slow 20-degree arc to the right.
9–10 seconds: subject settles; hold a clean final composition for the edit.
每個節拍都要做得到。一個主要動作加一個運鏡,通常比一串互不相關的變化清楚。讓時間區段加起來等於你選的時長,把重要的產品或臉部細節放在比較慢的節拍裡,並把最後半秒到一秒留給穩定的剪接點。
音訊可以共用同一個時鐘:
At 2.0 seconds, the first downbeat starts the hand movement.
At 6.0 seconds, the bass hit motivates the camera arc.
From 9.0 seconds, let the music tail continue under the held final frame.
想要更多鏡頭、光線和運鏡的詞彙,請看電影感 AI 影片提示詞指南。
六組 MiniMax H3 提示詞範本
把方括號裡的細節換掉,只附上列出的素材,並依你使用的介面調整標籤。這些範本是依據文件記載的 H3 輸入模式整理出來的結構化起點,不是宣稱得過獎的成績,也不保證輸出結果。
範本 1:帶動作與音樂參考的產品揭示
- 模式: 參考生成
- 素材: 一張乾淨的產品圖、一段運鏡影片,音樂參考選配
Create a 10-second 9:16 product reveal for a paid social placement.
Role map:
- Reference image 1 controls the product's exact design, proportions, packaging,
colors, materials, label placement, and wordmark.
- Reference video 1 controls camera path and acceleration only. Ignore its subject,
location, lighting, color grade, and audio.
- Reference audio 1 controls beat and edit timing only.
Scene: The product stands centered on a warm-violet studio plinth. Soft lilac key
light from camera left, restrained coral rim light, subtle atmospheric haze.
Timeline:
- 0–2 seconds: static wide reveal on the first soft beat.
- 2–7 seconds: transfer the smooth push-and-arc movement from reference video 1.
- 7–9 seconds: one narrow highlight travels across the product surface.
- 9–10 seconds: camera settles; hold the label front-facing and readable.
Preserve reference image 1 exactly: silhouette, cap, material finish, label layout,
brand colors, and wordmark. Keep one product only. No redesigned packaging,
duplicate objects, extra text, warped label, or abrupt camera shake.
這條提示詞之前的靜態圖準備工作,請用產品照轉影片流程。
範本 2:移轉角色動作而身分不跑掉
- 模式: 參考生成
- 素材: 一張角色圖、一段動作影片,環境圖選配
Create an 8-second 16:9 cinematic character performance.
Priority order:
1. Reference image 1 controls face, hair, body proportions, and wardrobe.
2. Reference video 1 controls body choreography and action timing only.
3. Reference image 2 controls the environment palette and architecture only.
4. This prompt controls framing and lighting.
The character from reference image 1 performs the complete movement from reference
video 1 in a moonlit station based on reference image 2. Medium full shot, camera
locked at chest height, soft directional light, realistic weight and foot contact.
Timeline:
- 0–1 seconds: character holds the starting pose.
- 1–7 seconds: perform the reference choreography once at its original tempo.
- 7–8 seconds: settle naturally and look toward camera left.
Preserve facial structure, hairstyle, coat shape, coat color, boots, and body
proportions from reference image 1 in every frame. Do not inherit the motion video's
performer, face, clothing, background, camera movement, or audio. No extra limbs,
sliding feet, costume changes, or cuts.
範本 3:帶聲音參考的對白表演
- 模式: 參考生成
- 素材: 一張角色圖、一段你有權使用的人聲或對白音訊
注意
只上傳你擁有、或已取得明確授權的聲音。模型接受音訊參考, 不代表你就有權模仿另一個人的 聲音。
Create a 9-second 16:9 single-character dialogue shot.
Role map:
- Reference image 1 controls the speaker's appearance, wardrobe, and room design.
- Reference audio 1 controls spoken timing, cadence, emotional progression, and pauses.
Shot: Medium close-up, eye-level, 50 mm cinematic framing. The speaker begins calm,
briefly smiles after the central pause, then finishes with quiet confidence. Natural
blinks and restrained hand movement. Camera remains static; soft room tone underneath.
Timeline:
- 0–1 seconds: silent eye contact and a small inhale.
- 1–8 seconds: performance follows reference audio 1 exactly in timing and pauses.
- 8–9 seconds: mouth closes, expression settles, hold for the edit.
Preserve face, hairstyle, skin tone, jacket, background layout, and lighting direction
from reference image 1. Keep one speaker. No camera move, cutaway, background speech,
new words, exaggerated gestures, or wardrobe changes.
範本 4:換掉來源影片的環境
- 模式: 參考生成
- 素材: 一段來源影片、一張環境圖,主體圖選配
Create a 12-second 16:9 environmental replacement based on the source clip.
Role map:
- Reference video 1 controls shot length, subject action, physical timing, camera path,
framing progression, and interaction with the ground.
- Reference image 1 controls the replacement environment, architecture, palette,
weather, and lighting mood.
- Reference image 2 controls the subject's face and wardrobe, if supplied.
Replace the source video's location with the environment from reference image 1.
Preserve the original action and camera movement continuously. Match subject lighting,
contact shadows, reflections, and atmospheric perspective to the new environment.
Timeline follows reference video 1. Maintain one continuous shot with the original
action beats and no added event.
Preserve the source subject's position, scale, motion, and ground contact. Preserve
identity and wardrobe from reference image 2 when present. Do not copy the original
background, signage, bystanders, color grade, or source audio. No cuts, teleporting,
floating feet, or changing architecture.
範本 5:對上節拍的品牌 UI 動態
- 模式: 參考生成
- 素材: 一張 UI 版面圖、一段動態風格影片,音訊參考選配
Create a 7-second 1:1 UI motion-design clip for a product announcement.
Priority order:
1. Reference image 1 controls exact layout, hierarchy, component positions, text,
logo, colors, corner shapes, and typography appearance.
2. Reference video 1 controls transition character and easing only.
3. Reference audio 1 controls the timing of three motion beats only.
Animate the interface from reference image 1 without redesigning it. Begin with the
main panel at rest. Use the restrained slide-and-scale transition quality from
reference video 1. Keep the camera orthographic and the background static.
Timeline:
- 0–1 seconds: complete layout at rest.
- 1–3 seconds: cards enter in sequence on beat one.
- 3–5 seconds: primary control changes state on beat two.
- 5–6 seconds: one subtle emphasis pulse on beat three.
- 6–7 seconds: complete interface holds sharp and readable.
Preserve every word, logo shape, component proportion, spacing relationship, and brand
color from reference image 1. No invented labels, misspelled text, extra panels,
perspective tilt, camera movement, glow overload, or elastic distortion.
範本 6:精準的首尾影格變化
- 模式: 首尾影格圖生影片
- 素材: 一張首影格圖和一張尾影格圖;不要附參考素材
Create a 10-second transformation from the supplied first frame to the supplied last
frame. Preserve the subject's identity and the camera's fixed position throughout.
Timeline:
- 0–2 seconds: hold the first-frame composition; only subtle ambient movement.
- 2–7 seconds: the scene transforms progressively from the center outward. Materials
change continuously with believable physical contact and no hard cut.
- 7–9 seconds: remaining details resolve into the last-frame design.
- 9–10 seconds: arrive exactly at the supplied last frame and hold it cleanly.
One continuous locked shot. Preserve subject scale, face, silhouette, horizon, lens,
and framing during the transition. No reference-style transfer, new characters,
camera movement, jump cut, flicker, or overshoot beyond the final composition.
不要把其他範本的參考圖片、影片或音訊附到這次請求上。官方 API 把首尾影格生成和參考生成視為互斥的模式。
一套實際可行的迭代順序
不要在一次結果不理想之後,就把身分、動作、運鏡、風格和音訊指示全部改掉。那樣你永遠學不到是哪一條指示起了作用。
分三輪迭代:
- 先鎖角色分工和不變項。 用最小的素材組,確認主體或產品仍然認得出來。
- 再鎖動作和時間軸。 加上動作來源或有時間點的節拍,風格先維持簡單。
- 最後加處理。 等前兩層穩定了,再放進環境、光線、調色、聲音提示和收尾細節。
結果走鐘的時候,先簡化,不要急著加更多否定語句。拿掉最不重要的那份參考、把時間軸縮短,或把某個來源的任務再收窄。用同一份檢查表比較幾個版本:身分、結構、動作、運鏡、聲音時間點,以及最後一格的穩定度。
從 OmniArt 開始
MiniMax H3 在 OmniArt 上開放給 Creator 以上的使用者。打開 MiniMax H3 影片生成頁,選擇 H3,再挑符合需求的模式:標準生成用 Video,首尾影格鏡頭用 Transition,混合圖片、影片和音訊條件則用 Reference。目前的產品介面支援 5–15 秒、1440p 的輸出。
在 Reference 模式下,最多可以附上 9 張圖片、3 段影片和 3 段音訊,總數不超過 12 份素材。參考音訊必須至少搭配一張圖片或一段影片。給每個附件指定一個角色,把對應的角色分工表和保留規則貼進提示詞,選好 5–15 秒的時長,然後生成。原生音訊是 H3 輸出的一部分,OmniArt 目前沒有另外開放音訊開關。
OmniArt 依成片每秒收 9 點數。參考影片每一秒(無條件進位)另收 9 點數;前 5 張參考圖免費,之後每多一張收 2 點數;參考音訊不另外收費。想知道 OmniArt 之外的官方 API 費率和成本試算,請看 MiniMax H3 價格與規格指南。
準備好開始創作了嗎?
用 AI 生成精彩內容