guide教程与操作指南17 分钟阅读

Seedance 2.5 参考素材提示词:给每个素材分配职责

Seedance 2.5 参考素材提示词的完整写法:给每张图片、每段视频和音频分配职责,把主体与素材绑定,再按场景挑选参考,让多主体镜头稳定生成。

OmniArt 团队
Seedance 2.5 参考素材提示词:给每个素材分配职责

Seedance 2.5 的参考素材提示词,和大多数创作者最先学会的文生视频写法完全不是一回事。模型在一次生成里最多可以接收 50 个参考素材——图片、视频和音频混在一起——而一旦超过两三个,难点就不再是描述场景了,而是告诉模型:哪个素材代表什么,以及每个瞬间该用到哪些素材。

这个转变值得直说。只有一张参考图时,模型可以猜出你的意图;有十二张时,猜测就会带来五官融合、服装互换、道具重复,以及从「只想取一张脸」的图片里渗进来的背景。解决办法不是把场景描述写得更长,而是在提示词开头加一小段明确的说明,给每个素材分配一份工作和一组排除项。

Seedance 2.5 已经在 OmniArt 的视频创作工作台上线,你可以边读边用自己的素材跟着做。这份指南会讲清核心提示词公式、素材数量限制、职责定义写法,以及一次要承载多个主体时的四步流程。

从核心提示词公式开始

在参考素材登场之前,Seedance 2.5 遵循的是一个简单灵活的公式。前两项之后的所有元素都是可选的——能改变镜头的就写,不能的就删。

主体 + 动作或事件 + 场景与环境 + 视觉风格 + 运镜或剪辑 + 音频

写成模板是这样:

<Subject> performs <primary action or event> in <scene and environment>.
The visuals feature <visual style>.
Use <shot size, camera angle, camera movement, or cuts>.
Audio includes <dialogue, ambience, sound effects, or music>.

填成一个真实镜头:

A bicycle mechanic trues a wheel at a workbench in a narrow street-front workshop, then lifts the wheel onto a wall hook.
Late afternoon light comes through the roll-up door. Metal surfaces are worn and slightly oily, and the floor is swept.
Begin with a medium shot of the truing stand, push in slowly toward the spoke nipples, then cut to a frontal view of the wall.
Retain the ping of spoke tension, the click of the truing gauge, and quiet street ambience.

分辨率、时长这类生成参数不属于提示词——在生成页面上设置就好。提示词里应该只留下模型必须去理解的内容。

动笔之前先搞清素材预算

Seedance 2.5 最多合计接收 50 个参考素材,分三类。每一类都有一个硬性输入上限,和一个单独的推荐范围。推荐范围不是封顶值,而是生成结果保持稳定的区间。

素材类型输入上限推荐范围
图片最多 30 张,每张不超过 4K主体参考图里 1–8 个不同主体
视频最多 10 个,合计不超过 30 秒1–5 个不同主体,每个主体视频 5–10 秒
音频最多 10 段,合计不超过 30 秒只放与任务相关的对白、人声、环境声或音乐
视频编辑一个源视频加参考图片源视频不超过 20 秒,1–5 张参考图

你可以越过推荐范围——主体图里放 9–12 个主体、主体音频或视频里放 6–10 个主体、一次编辑用 6–8 张参考图。数量越多,稳定性通常越低,所以任何超出推荐范围的做法,都应该当作一次有意为之的实验,而不是默认设置。

有一条打包规则的重要性超出大多数人的预期:如果主体超过五个、而且还需要多个视角,就把每个视角单独放进一张图片。按视角拆分的独立图片,通常比把几个视角拼进一张图更稳定。一张四格拼图在模型看来就是「一张里面有四样东西的图片」,而这正是你想消除的那种歧义。

提示

上传之前先做减法。十个精准、单一职责的素材,胜过二十个互相重叠的素材,因为每多一个素材,模型就多一件要消歧的事。

给每个素材定义职责——以及它的排除项

这是这份指南后半部分的地基。上传完素材之后,明确说出每个素材贡献什么,以及有哪些东西绝对不能被带过来。素材映射必须写在提示词文本里;不要指望图片上烧录的标签,也不要让模型自己去猜某个文件代表哪个人物或哪件道具。

写法是这样:

@Image 1 defines <subject>'s <appearance, clothing, structure, or material>.
@Video 1 defines <motion, camera movement, or pacing>.
@Audio 1 defines <character or sound type>'s <voice, dialogue, ambience, or music>.

<Subject> completes <primary action or event> in <scene>.
The visuals feature <visual style>, with <camera treatment>.

填好之后,排除项会实实在在地起作用:

@Image 1 defines the mechanic's facial features, hairstyle, and navy work apron. Do not use the image background.
@Image 2 defines the workshop's bench layout, tool wall, and roll-up door light. Do not use the people in the image.
@Video 1 defines the pacing of spinning the wheel, pausing, and tightening a spoke. Do not use the person's identity, clothing, or scene from the video.

The mechanic trues a wheel at the workbench, then lifts it onto the wall hook.
Begin with a medium shot of the truing stand and push in slowly toward the spoke nipples. Retain spoke pings and quiet street ambience.

那几行排除项不是凑字数。一张人像参考图必然带着背景、一个打光方向,而且往往还有别的人。如果不写上「不要使用图片背景」,这些属性对模型来说就是可用的,而它们迟早会出现在画面的某个地方。

把多个视角绑定到同一个对象

当好几张图片从不同角度展示同一个人或同一件产品时,要明确说出来,并且说清最终画面里应该存在几个这样的对象:

@Image 1 defines the front view of the same steel touring frame.
@Image 2 defines the left-side structure of the same steel touring frame.
@Image 3 defines the right-side structure of the same steel touring frame.
@Image 4 defines the rear structure of the same steel touring frame.
All four images define one steel touring frame. The output must contain only one frame throughout.

没有最后那句话,「同一个对象的四张图」也完全可以被理解成「四个对象」。这句数量断言是整条提示词里最便宜的保险。

多参考素材的四步流程

一旦素材超过几个,就按顺序走完下面四步。目标是帮模型为当前场景挑出正确的素材,而不是让所有素材同时出现。

第一步:逐个命名并映射每个主体

把每个人物、产品和道具都绑定到属于它自己的素材上,一行写一个:

<Character A> corresponds to @Image 1. Use only the appearance, hairstyle, and clothing.
<Character B> corresponds to @Image 2. Use only the appearance, hairstyle, and clothing.
<Prop A> corresponds to @Image 3. Use only the structure, material, and color.
<Scene A> references @Image 4. Use only the spatial layout, architecture, and lighting. Do not use the people in the image.

要避开的反面写法是:永远不要写「@Image 1 到 @Image 4 分别定义四个角色」。这句话只告诉模型有四个角色和四张图片,却从来没说哪张对应哪个——于是它可以随意配对。

第二步:按类型给素材分组

映射关系都写好之后,用小标题把它们组织起来,让关系一眼可读:

[Characters]
<Mechanic> corresponds to @Image 1. Use only the appearance, hairstyle, and clothing.
<Apprentice> corresponds to @Image 2. Use only the appearance, hairstyle, and clothing.
Do not interchange the two characters' appearances, clothing, actions, positions, or dialogue.

[Props]
<Truing Stand> corresponds to @Image 3 and belongs only to <Mechanic>.
<Parts Tray> corresponds to @Image 4 and belongs only to <Apprentice>.

[Scenes]
<Workshop> references @Image 5. Use only the space, materials, and lighting.
<Courtyard> references @Image 6. Use only the space, materials, and lighting.

[Motion and Audio]
@Video 1 defines the motion of <Mechanic> tightening a spoke. Do not use the person or scene from the video.
@Audio 1 defines <Apprentice>'s voice and specified dialogue.

道具归属那几行和角色行同样重要。「只属于某人」这句话,正是阻止一件工具在两个人之间跨镜头瞬移的关键。

第三步:为重要主体写一份档案

当一个角色横跨多个场景、要用到好几个素材时,把它们集中写成一整块,而不是散落在各处:

[Subject Profile: Mechanic]
Appearance and clothing: @Image 1.
Fixed prop: <Truing Stand> from @Image 3.
Locations: <Workshop> and <Courtyard>.
Motion references: the spoke-tightening motion from @Video 1 and the wheel-lifting motion from @Video 2.
Do not use: the apprentice's clothing. Do not give this character <Parts Tray>.

第四步:按场景挑选参考素材

最后,告诉模型每个场景使用素材里的哪一个子集,以及这个场景结束时画面应该是什么样:

Scene 1 | Truing the wheel in the workshop
Use: <Mechanic>, <Truing Stand>, <Workshop>, and the spoke-tightening motion from @Video 1.
Event: <Mechanic> spins the wheel in <Truing Stand> and tightens one spoke.
End state: <Mechanic> stands on the inner side of the bench. <Truing Stand> stays at the left of the frame.

Scene 2 | Handoff in the courtyard
Use: <Apprentice>, <Parts Tray>, and <Courtyard>.
Event: <Apprentice> carries <Parts Tray> to the courtyard bench and sets it down.
End state: <Parts Tray> rests on the bench. No other character enters the courtyard.

按场景挑选素材,才能把一堆参考变成一份分镜表。凡是你没在某个场景里点到名的素材,模型就知道要把它排除在外。

音频与文字的专用语法

提示词完全可以用自然语言写,但当你需要把音乐、音效、对白和字幕严格区分开时,Seedance 2.5 认得四种括号:

内容语法示例
音乐()(Soft, rhythmic piano music plays in the background)
音效<><A bell rings in the distance>
对白{}{Hello, welcome back.}
字幕【】【Chapter One: Departure】

强化对白语言

当对白不是中文时,先在台词前面点明语言——例如 The girl says softly in Japanese: {もう大丈夫です}。如果写下来的文本是英文、但模型却用别的语言念出来,或者你需要某种特定的地区口音,就用这个公式:

对白语言 + 地区变体或口音 + 表达方式 + 说话人 + {dialogue}

Dialogue language: American English. The girl says in natural, conversational American English: {I thought you weren't coming.}

Dialogue language: authentic Los Angeles English. The young man says in natural Los Angeles vernacular: {No way, you actually made it.}

参考素材已经定义过的,不要再复述一遍

一种常见的过度书写习惯是:上传了一段运动参考视频,然后又在正文里把同一段运动再描述一次。当参考视频已经准确定义了运动、镜头和顺序时,只需要说明要继承哪些属性就够了。重复描述运动会和参考素材本身冲突,模型就得在同一条指令的两个版本之间做取舍。

例外是白模视频。白模主要提供运动和空间结构——路径、调度、运镜、剪辑点——所以提示词仍然要定义主体、场景、动作和视觉风格。时间节奏从视频继承,画面世界由图片和文字提供。如果白模参考是你的主要工作流,Seedance 2.5 分镜与白模指南会更深入地讲粗颗粒白模与精细白模的区别。

提交前过一遍清单

按下生成之前,对照这几个问题读一遍你的提示词:

  • 提示词是否清楚说明了主体和主要动作或事件?
  • 每个参考素材是否都写明了「用什么」和「不用什么」?
  • 每个不同的角色、产品和道具是否都被命名并绑定到了参考素材?
  • 参考素材是否按场景挑选,而不是被要求全部同时出现?
  • 角色数量、服装、道具归属和空间关系是否保持一致?
  • 抽象的情绪和摄影术语,是否都配上了看得见或听得见的具体线索?

警告

参考素材提示词提高的是结果忠实的概率,而不是保证。字幕、公式、招牌文字、产品参数,或者必须精确的帧级时间点,都要靠事先准备的参考素材、生成和后期共同完成。

在 OmniArt 上开始使用

打开 OmniArt 的视频创作工作台,选择 Seedance 2.5,按这个顺序写下你的第一条多参考提示词:

  1. 只上传那些携带了别的素材无法提供的信息的素材。
  2. 每个主体写一行映射——名字、素材,以及要从中取用什么。
  3. 凡是背景、人物或构图可能渗出来的地方,都给那一行加上排除项。
  4. [Characters][Props][Scenes][Motion and Audio] 把映射分组。
  5. 给任何横跨多个场景的角色写一份主体档案。
  6. 逐个列出每个场景所用的素材子集、发生的事件和结束状态。
  7. 用同一条提示词生成两次,凡是两次结果不一致的地方,去改映射关系,而不是改形容词。

关于这次发布具体更新了什么、新模式各自适合什么,可以读 Seedance 2.5 上线盘点。想把同样的参考素材纪律带进更长的片子,可以继续看 30 秒 Seedance 2.5 视频提示词写法;如果需要贯穿整个项目的身份一致性,可以参考我们的 AI 视频角色一致性指南

准备好创作了吗?

开始用 AI 生成精彩内容

免费开始