guide教程与操作指南21 分钟阅读

MiniMax H3 提示词指南:掌控每一种参考输入

学习 MiniMax H3 在图片、视频、音频与文本参考下的提示词写法。为每个素材绑定职责,避免冲突,并指挥一段有时间线的 AI 视频镜头。

OmniArt 团队
MiniMax H3 提示词指南:掌控每一种参考输入

MiniMax H3 提示词真正有用的思路,不是「写得更细」,而是「给每个输入只分配一件事」。H3 可以把文本、图片、视频和音频作为统一参考集来接收,因此创作者可以从静帧提取产品身份、从片段提取运动、从音频提取节奏,再从提示词给出镜头调度。只有当模型分得清哪一个来源控制结果的哪一部分时,这种灵活性才真正有用。

这份 MiniMax H3 提示词指南会把上述原则落成可复用的写法。你会学到如何为每个素材绑定职责、防止参考之间互相矛盾、撰写可检查的保留规则,以及如何把 4–15 秒的生成拆成带时间码的节拍。文末六个模板分别覆盖产品视频、角色动作、对白、视频剪辑、UI 动效,以及首尾帧转场。

说明

这是一份以文档为依据的提示词指南,不是 OmniArt 实测 H3 生成结果的报告。内容基于 MiniMax 当前的 H3 生成指南官方 API 参考。 截至 2026 年 7 月 31 日,H3 尚未在 OmniArt 的权威模型目录中被确认可用。界面标签与参考语法 在 MiniMax 自有产品面与下游托管方之间也可能不同。

MiniMax H3、Hailuo 3.0 还是 Hailuo 03?

MiniMax 当前文档将模型称为 MiniMax H3,官方 API 模型标识符为 MiniMax-H3。你也可能在发布报道与第三方目录中看到 Hailuo 3.0Hailuo 03Hailuo-3.0,因为 Hailuo 是 MiniMax 的视频产品品牌。

请使用你面前那个产品面所要求的标签。就发布时文档中的官方 MiniMax V2 API 而言,这意味着 MiniMax-H3;不要假定市场别名同样是有效的官方 API 标识符。本指南用 MiniMax H3 指代模型,用 参考生成 指代可同时组合图片、视频与音频输入的模式。「Omni 参考」是对这种混合媒体工作流的实用描述,但官方 API 称之为 reference-to-video。

先搞清 H3 的三种生成模式

选对模式,再写提示词。

模式输入适用场景
文生视频必需的文本提示词场景可以仅靠描述创建
首/尾帧图生视频必需的提示词,外加首帧、尾帧,或两者都有需要精确的开场或收尾构图
参考生成必需的提示词,外加参考图片、视频或音频你需要从上传素材迁移身份、设计、运动、镜头、人声、节奏或风格

当前官方规格列出 2K 输出,以及 4 到 15 秒的整数时长。每次请求都需要非空文本提示词,上限 7,000 个字符。参考请求最多可包含 9 张图片、3 个视频和 3 个音频片段,素材总数不超过 12 个。参考视频合计时长上限为 15 秒,参考音频片段同样如此。音频参考不能单独提交;它必须至少伴随一张图片或一个视频。

文生视频需要明确的宽高比。参考生成可以使用明确比例,也可以让 H3 自适应选择;首/尾帧生成则跟随上传图片。文档中的比例选项为 21:9、16:9、4:3、1:1、3:4 与 9:16。

有一条硬性工作流边界:首/尾帧模式与参考生成不能在同一次 API 请求中混用。如果精确的边界帧更重要,用首/尾帧模式。如果身份、运动、风格或音频迁移更重要,用参考生成,并在文本中描述开场构图。

把 H3 提示词建成职责地图

H3 的 API 以带类型的条目接收媒体,角色如 reference_imagereference_videoreference_audio。这些技术角色标明媒体类型;你的提示词应再进一步,说明每个素材的创意职责。

按这个顺序写:

DELIVERABLE
What the final clip is for, its duration, and its format.

ROLE MAP
Reference image 1 = subject identity and immutable appearance.
Reference image 2 = environment or visual style only.
Reference video 1 = body motion and timing only.
Reference audio 1 = rhythm and edit timing only.

SHOT
Subject, action, environment, lighting, and camera direction.

TIMELINE
Time-coded beats that add up to the selected duration.

PRESERVE
Features that must remain unchanged from named references.

EXCLUDE
A short list of the most damaging unwanted changes.

上面的编号标签是你简报中的语义标签,并不承诺每个 H3 界面都会暴露字面的 Reference image 1 句柄。在官方 API 中,文件是 content 数组里的独立条目。在可视化界面中,使用该产品面提供的附件名或提及语法,并保持同一套「一个素材一件事」的逻辑。

给每个素材只分配一个主责

参考素材包含的信息,通常比你真正想迁移的更多。一支舞蹈片段里可能有表演者、服装、场地、镜头、光线、动作,或许还有音乐。如果你让 H3「跟随视频」,这些特征会与你的角色图互相抢控制权。

改为窄绑定:

  • 图片参考: 身份、面部、服装、产品几何、logo、色板或环境设计。
  • 视频参考: 肢体动作、镜头路径、动作时序、物理交互或剪辑节奏。
  • 音频参考: 声线特征、语速节奏、音乐节拍、环境音或事件提示。
  • 文本提示词: 最终场景、迁移规则、时间线、优先级与保留指令。

当一个素材必须承担两件事时,把两者都写出来,并把边界说清楚:「Reference video 1 控制镜头路径与动作时序;忽略其表演者、服装、场地、调色与音频。」

防止参考素材互相打架

参考冲突通常先是指令问题,然后才是模型问题。两张图可能显示不同外套,动作片段可能使用不同体型,或音频节拍暗示的切点与你要求的不间断运镜冲突。

只要输入有重叠,就写明优先级:

Priority order:
1. Reference image 1 controls character identity and wardrobe.
2. Reference video 1 controls movement and timing only.
3. Reference image 2 controls lighting palette and set design only.
4. The text prompt controls camera framing and final composition.

然后主动消解矛盾,而不是指望 H3 猜出你的偏好。若身份图是披发,动作视频却戴帽子,就写明最终角色是保留头发还是戴帽子。若产品参考有可读标签,而风格图偏绘画感,就说明风格化只作用于环境,包装保持写实。

三种习惯能减少冲突:

  1. 用最小可用参考集。 输入越多,可能分歧越多。只有当某文件贡献了提示词难以可靠表达的信息时,才加入它。
  2. 把内容与风格分开。 写明哪个参考拥有主体,哪个拥有处理方式。
  3. 只选一个镜头权威。 要么从视频迁移镜头,要么用文本指挥镜头。若两者都需要,写明视频提供时序,文本覆盖构图。

若需要按症状排查的更广工作流,迭代时把 7 个 AI 视频提示词修复 放在本指南旁对照使用。

写可检查的保留指令

「保持一致」太模糊。保留块应点名你能在任意一帧里检查的可见属性。

对角色:

Preserve from reference image 1 throughout every frame: facial structure,
eye color, hairstyle and length, jacket cut, jacket color, and body proportions.
Do not replace the performer with the person from reference video 1.

对产品:

Preserve from reference image 1: bottle silhouette, cap shape, label placement,
navy wordmark, coral seal, material finish, and relative proportions.
No duplicate product, redesigned packaging, extra text, or changing logo.

先写正向锚点;排除项是护栏,不是整段提示词。否定列表保持简短且具体。十条含糊禁令,可能冲淡真正决定成片是否可用的四条不变量。

如果同一角色会出现在多条分别生成的片段中,在每条提示词中保持身份与服装块完全一致。一致角色工作流 说明了如何在更长序列中构建可复用的参考包。

指挥时间,而不只是内容

H3 允许 4 到 15 秒的片段,但一大段段落并不会告诉模型每个事件何时发生。时间线把提示词变成紧凑的镜头计划。

对于 10 秒片段:

0–2 seconds: locked medium-wide establishing shot; subject holds still.
2–6 seconds: subject performs the referenced action at the source tempo.
6–9 seconds: camera makes one slow 20-degree arc to the right.
9–10 seconds: subject settles; hold a clean final composition for the edit.

让每个节拍都可实现。一个主动作加一次运镜,通常比一串无关变形更清晰。让时间区间加总等于所选时长,把重要的产品或面部细节放在较慢的节拍,并把最后半秒或一秒留给稳定剪辑点。

音频可以共享同一时钟:

At 2.0 seconds, the first downbeat starts the hand movement.
At 6.0 seconds, the bass hit motivates the camera arc.
From 9.0 seconds, let the music tail continue under the held final frame.

需要更多镜头、光线与运镜词汇时,见 电影感 AI 视频提示词指南

六个 MiniMax H3 提示词模板

替换括号内细节,只挂载列出的素材,并把标签适配到你使用的界面。这些模板是基于已记录 H3 输入模式的结构化起点;它们不是声称的基准胜者,也不保证输出结果。

模板 1:结合运动与音乐参考的产品揭示

  • 模式: 参考生成
  • 素材: 一张干净产品图、一段运镜视频、可选音乐参考
Create a 10-second 9:16 product reveal for a paid social placement.

Role map:
- Reference image 1 controls the product's exact design, proportions, packaging,
  colors, materials, label placement, and wordmark.
- Reference video 1 controls camera path and acceleration only. Ignore its subject,
  location, lighting, color grade, and audio.
- Reference audio 1 controls beat and edit timing only.

Scene: The product stands centered on a warm-violet studio plinth. Soft lilac key
light from camera left, restrained coral rim light, subtle atmospheric haze.

Timeline:
- 0–2 seconds: static wide reveal on the first soft beat.
- 2–7 seconds: transfer the smooth push-and-arc movement from reference video 1.
- 7–9 seconds: one narrow highlight travels across the product surface.
- 9–10 seconds: camera settles; hold the label front-facing and readable.

Preserve reference image 1 exactly: silhouette, cap, material finish, label layout,
brand colors, and wordmark. Keep one product only. No redesigned packaging,
duplicate objects, extra text, warped label, or abrupt camera shake.

在这条提示词之前的静帧准备,见 照片转产品视频工作流

模板 2:角色动作迁移且不漂移身份

  • 模式: 参考生成
  • 素材: 一张角色图、一段动作视频、可选环境图
Create an 8-second 16:9 cinematic character performance.

Priority order:
1. Reference image 1 controls face, hair, body proportions, and wardrobe.
2. Reference video 1 controls body choreography and action timing only.
3. Reference image 2 controls the environment palette and architecture only.
4. This prompt controls framing and lighting.

The character from reference image 1 performs the complete movement from reference
video 1 in a moonlit station based on reference image 2. Medium full shot, camera
locked at chest height, soft directional light, realistic weight and foot contact.

Timeline:
- 0–1 seconds: character holds the starting pose.
- 1–7 seconds: perform the reference choreography once at its original tempo.
- 7–8 seconds: settle naturally and look toward camera left.

Preserve facial structure, hairstyle, coat shape, coat color, boots, and body
proportions from reference image 1 in every frame. Do not inherit the motion video's
performer, face, clothing, background, camera movement, or audio. No extra limbs,
sliding feet, costume changes, or cuts.

模板 3:带声音参考的对白表演

  • 模式: 参考生成
  • 素材: 一张角色图、一段已获授权的人声或对白音频片段

警告

只上传你拥有或明确获得使用许可的声音。模型接受参考音频,并不赋予你模仿他人声音的权利。

Create a 9-second 16:9 single-character dialogue shot.

Role map:
- Reference image 1 controls the speaker's appearance, wardrobe, and room design.
- Reference audio 1 controls spoken timing, cadence, emotional progression, and pauses.

Shot: Medium close-up, eye-level, 50 mm cinematic framing. The speaker begins calm,
briefly smiles after the central pause, then finishes with quiet confidence. Natural
blinks and restrained hand movement. Camera remains static; soft room tone underneath.

Timeline:
- 0–1 seconds: silent eye contact and a small inhale.
- 1–8 seconds: performance follows reference audio 1 exactly in timing and pauses.
- 8–9 seconds: mouth closes, expression settles, hold for the edit.

Preserve face, hairstyle, skin tone, jacket, background layout, and lighting direction
from reference image 1. Keep one speaker. No camera move, cutaway, background speech,
new words, exaggerated gestures, or wardrobe changes.

模板 4:替换源视频环境

  • 模式: 参考生成
  • 素材: 一段源视频、一张环境图、可选主体图
Create a 12-second 16:9 environmental replacement based on the source clip.

Role map:
- Reference video 1 controls shot length, subject action, physical timing, camera path,
  framing progression, and interaction with the ground.
- Reference image 1 controls the replacement environment, architecture, palette,
  weather, and lighting mood.
- Reference image 2 controls the subject's face and wardrobe, if supplied.

Replace the source video's location with the environment from reference image 1.
Preserve the original action and camera movement continuously. Match subject lighting,
contact shadows, reflections, and atmospheric perspective to the new environment.

Timeline follows reference video 1. Maintain one continuous shot with the original
action beats and no added event.

Preserve the source subject's position, scale, motion, and ground contact. Preserve
identity and wardrobe from reference image 2 when present. Do not copy the original
background, signage, bystanders, color grade, or source audio. No cuts, teleporting,
floating feet, or changing architecture.

模板 5:与节拍同步的品牌 UI 动效

  • 模式: 参考生成
  • 素材: 一张 UI 布局图、一段动效风格视频、可选音频参考
Create a 7-second 1:1 UI motion-design clip for a product announcement.

Priority order:
1. Reference image 1 controls exact layout, hierarchy, component positions, text,
   logo, colors, corner shapes, and typography appearance.
2. Reference video 1 controls transition character and easing only.
3. Reference audio 1 controls the timing of three motion beats only.

Animate the interface from reference image 1 without redesigning it. Begin with the
main panel at rest. Use the restrained slide-and-scale transition quality from
reference video 1. Keep the camera orthographic and the background static.

Timeline:
- 0–1 seconds: complete layout at rest.
- 1–3 seconds: cards enter in sequence on beat one.
- 3–5 seconds: primary control changes state on beat two.
- 5–6 seconds: one subtle emphasis pulse on beat three.
- 6–7 seconds: complete interface holds sharp and readable.

Preserve every word, logo shape, component proportion, spacing relationship, and brand
color from reference image 1. No invented labels, misspelled text, extra panels,
perspective tilt, camera movement, glow overload, or elastic distortion.

模板 6:精确的首尾帧变换

  • 模式: 首/尾帧图生视频
  • 素材: 一张首帧图与一张尾帧图;不要附带参考媒体
Create a 10-second transformation from the supplied first frame to the supplied last
frame. Preserve the subject's identity and the camera's fixed position throughout.

Timeline:
- 0–2 seconds: hold the first-frame composition; only subtle ambient movement.
- 2–7 seconds: the scene transforms progressively from the center outward. Materials
  change continuously with believable physical contact and no hard cut.
- 7–9 seconds: remaining details resolve into the last-frame design.
- 9–10 seconds: arrive exactly at the supplied last frame and hold it cleanly.

One continuous locked shot. Preserve subject scale, face, silhouette, horizon, lens,
and framing during the transition. No reference-style transfer, new characters,
camera movement, jump cut, flicker, or overshoot beyond the final composition.

不要把其他模板的参考图片、视频或音频挂到这次请求上。官方 API 将首/尾帧生成与参考生成视为互斥模式。

实用的迭代顺序

不要在一次令人失望的成片后,同时改身份、运动、镜头、风格与音频指令。那样你无法学到哪条指令真正有效。

分三轮迭代:

  1. 锁定职责与不变量。 使用最小参考集,确认主体或产品始终可识别。
  2. 锁定动作与时间线。 加入运动来源或带时间的节拍,风格保持简单。
  3. 再加处理。 等前两层稳定后,再引入环境、光线、调色、音频提示与收尾细节。

结果发生漂移时,先简化,再加更多否定语。移除最不重要的参考,缩短时间线,或把某一来源的职责收窄。用同一清单比较多个变体:身份、几何、运动、镜头、音频时序,以及尾帧稳定性。

从 OmniArt 开始

MiniMax H3 目前尚未确认可在 OmniArt 内选择,因此本文不会链到 H3 创建路由,也不会声称存在 OmniArt 的 H3 工作流。查看 当前 OmniArt 视频模型阵容 了解今天可用的模型。

职责地图方法仍然能很好地迁移到受支持的参考驱动视频模型:在打开工作区前,准备一张身份图、一个运动来源,以及一段简短的保留块。在创意简报稳定前保持提示词模型中立,再把附件标签与功能专用语法适配到你实际选择的模型。这样你现在就有一份可复用的方向表,若模型之后加入 OmniArt,也能直接得到干净的 H3 提示词。

关于官方 API 费率、参考视频计费与成本算例,见 MiniMax H3 价格与规格指南。该文使用 MiniMax 自己的定价页,而非转售商费率。

准备好创作了吗?

开始用 AI 生成精彩内容

免费开始