Guide教學與操作指南閱讀時間 17 分鐘

Seedance 2.5 參考素材提示詞:給每一份素材一個職責

Seedance 2.5 參考素材提示詞逐步教學:替每張圖片、每段影片和每段音訊指派職責,把主體綁到素材上,再依場景挑選參考素材。

OmniArt 團隊
Seedance 2.5 參考素材提示詞:給每一份素材一個職責

Seedance 2.5 的參考素材提示詞,跟大多數創作者最早學的文生影片提示詞運作方式不一樣。這個模型一次生成吃得下最多 50 份參考輸入——圖片、影片和音訊加起來——而一旦你超過兩三份,難的就不再是描述一個場景了。難的是告訴模型哪一份素材代表什麼,以及哪幾份屬於影片裡的哪一個時刻。

這個轉變值得直說。只有一張參考圖時,模型可以猜你的意圖。有十二張時,猜出來的是混在一起的臉、換錯的衣服、複製出來的道具,還有從你只想拿臉的那幾張圖裡滲進來的背景。解法不是把場景描述寫得更長。而是在提示詞開頭放一小段明確的區塊,替每一份素材指派一個職責和一份排除清單。

Seedance 2.5 在 OmniArt 的影片創作工作台裡就有,所以你可以一邊讀一邊拿自己的參考素材跟著做。這份指南涵蓋核心提示詞公式、素材上限、職責定義的寫法,還有一次要扛很多主體時的四步流程。

從核心提示詞公式開始

在參考素材進場之前,Seedance 2.5 就對一條簡單、有彈性的公式反應良好。前兩項之後的每一個元素都是選填的——會改變這顆鏡頭的就寫進去,不會的就丟掉。

主體 + 動作或事件 + 場景與環境 + 視覺風格 + 運鏡或剪接 + 音訊

寫成範本:

<Subject> performs <primary action or event> in <scene and environment>.
The visuals feature <visual style>.
Use <shot size, camera angle, camera movement, or cuts>.
Audio includes <dialogue, ambience, sound effects, or music>.

填成一顆真的鏡頭:

A bicycle mechanic trues a wheel at a workbench in a narrow street-front workshop, then lifts the wheel onto a wall hook.
Late afternoon light comes through the roll-up door. Metal surfaces are worn and slightly oily, and the floor is swept.
Begin with a medium shot of the truing stand, push in slowly toward the spoke nipples, then cut to a frontal view of the wall.
Retain the ping of spoke tension, the click of the truing gauge, and quiet street ambience.

解析度和長度這類生成參數不該寫進提示詞裡——那些在生成頁面上設定。提示詞裡的每一樣東西,都應該是模型必須去詮釋的內容。

動筆之前先搞清楚你的素材預算

Seedance 2.5 把三種類型的參考輸入合起來,最多 50 份。每一種類型都有一個硬性的輸入上限,還有另外一個建議範圍。建議範圍不是上限——它是生成維持穩定的那個區間。

素材類型輸入上限建議範圍
圖片最多 30 張,每張不超過 4K跨主體參考圖 1–8 個不同主體
影片最多 10 段,合計長度不超過 30 秒1–5 個不同主體,每個主體影片 5–10 秒
音訊最多 10 段,合計長度不超過 30 秒只放跟這個任務相關的對白、人聲、環境音或音樂
影片編輯一段來源影片加參考圖來源影片 20 秒以內,1–5 張參考圖

你可以超過建議值——主體圖片放 9–12 個主體、主體音訊或影片放 6–10 個主體、一次編輯放 6–8 張參考圖。數量一多,穩定度通常就會下降,所以任何超過建議範圍的做法,都該當成一次刻意的實驗,而不是預設值。

有一條打包規則比大家想的更重要:如果超過五個主體還各自需要多個視角,就讓每一個視角自己一張圖。獨立的視角圖通常比把好幾個視角拼成一張拼版更穩定。一張四格拼版對模型來說就是「一張裡面有四樣東西的圖」,而那正是你想消掉的那種模糊。

小技巧

上傳之前先修剪。十份緊湊、單一用途的素材,表現勝過二十份互相重疊的,因為每多一份素材,就是模型必須多消歧一次的東西。

定義每一份素材的職責——還有它的排除項

這是這份指南接下來全部建立在上面的技巧。上傳之後,明確講出每一份素材貢獻什麼,也講出什麼絕對不能被帶過去。素材對應關係要寫在提示詞文字裡;不要靠燒在圖片上的標籤,也不要讓模型自己去推斷某個檔案代表哪一個人或哪一件道具。

寫法:

@Image 1 defines <subject>'s <appearance, clothing, structure, or material>.
@Video 1 defines <motion, camera movement, or pacing>.
@Audio 1 defines <character or sound type>'s <voice, dialogue, ambience, or music>.

<Subject> completes <primary action or event> in <scene>.
The visuals feature <visual style>, with <camera treatment>.

填好之後,排除項在真的幹活:

@Image 1 defines the mechanic's facial features, hairstyle, and navy work apron. Do not use the image background.
@Image 2 defines the workshop's bench layout, tool wall, and roll-up door light. Do not use the people in the image.
@Video 1 defines the pacing of spinning the wheel, pausing, and tightening a spoke. Do not use the person's identity, clothing, or scene from the video.

The mechanic trues a wheel at the workbench, then lifts it onto the wall hook.
Begin with a medium shot of the truing stand and push in slowly toward the spoke nipples. Retain spoke pings and quiet street ambience.

那幾行排除項不是填充物。一張人像參考圖會帶著一個背景、一個打光方向,而且常常還帶著別人。少了 "do not use the image background,",那些屬性對模型來說就是可以拿的,而且它們一定會在某個地方冒出來。

把多個視角綁到同一個物件上

當好幾張圖片從不同角度拍的是同一個人或同一件產品時,要明確講出來,而且要講清楚輸出裡該存在幾個這個物件:

@Image 1 defines the front view of the same steel touring frame.
@Image 2 defines the left-side structure of the same steel touring frame.
@Image 3 defines the right-side structure of the same steel touring frame.
@Image 4 defines the rear structure of the same steel touring frame.
All four images define one steel touring frame. The output must contain only one frame throughout.

少了最後那句話,「四張同一個物件的圖」也可以合理地被讀成「四個物件」。宣告數量是整句提示詞裡最便宜的保險。

四步式的多參考素材流程

一旦素材超過幾份,就照順序走完這四個步驟。目標是幫模型替當下這個場景挑出對的素材——而不是讓所有東西一次全部出現。

步驟 1:把每一個主體個別點名並對應

把每一個人、每一件產品、每一件道具都綁到它自己的素材上,一個一行:

<Character A> corresponds to @Image 1. Use only the appearance, hairstyle, and clothing.
<Character B> corresponds to @Image 2. Use only the appearance, hairstyle, and clothing.
<Prop A> corresponds to @Image 3. Use only the structure, material, and color.
<Scene A> references @Image 4. Use only the spatial layout, architecture, and lighting. Do not use the people in the image.

要避開的反面寫法:絕對不要寫 "@Images 1 through 4 define four characters respectively."。那句話告訴模型有四個角色和四張圖,卻從來沒說哪張配哪個——所以它可以隨自己高興去配。

步驟 2:把素材按類型分組

每一條對應都寫好之後,用標題把它們組織起來,讓關係一眼就讀得出來:

[Characters]
<Mechanic> corresponds to @Image 1. Use only the appearance, hairstyle, and clothing.
<Apprentice> corresponds to @Image 2. Use only the appearance, hairstyle, and clothing.
Do not interchange the two characters' appearances, clothing, actions, positions, or dialogue.

[Props]
<Truing Stand> corresponds to @Image 3 and belongs only to <Mechanic>.
<Parts Tray> corresponds to @Image 4 and belongs only to <Apprentice>.

[Scenes]
<Workshop> references @Image 5. Use only the space, materials, and lighting.
<Courtyard> references @Image 6. Use only the space, materials, and lighting.

[Motion and Audio]
@Video 1 defines the motion of <Mechanic> tightening a spoke. Do not use the person or scene from the video.
@Audio 1 defines <Apprentice>'s voice and specified dialogue.

道具歸屬那幾行,跟角色那幾行一樣重要。"Belongs only to" 就是那句擋住一支工具在兩個人之間、在剪點之間瞬移的話。

步驟 3:替重要的主體寫一份檔案

當一個角色跨好幾個場景、要用到好幾份素材時,把它們收進同一個區塊,不要讓它們散落各處:

[Subject Profile: Mechanic]
Appearance and clothing: @Image 1.
Fixed prop: <Truing Stand> from @Image 3.
Locations: <Workshop> and <Courtyard>.
Motion references: the spoke-tightening motion from @Video 1 and the wheel-lifting motion from @Video 2.
Do not use: the apprentice's clothing. Do not give this character <Parts Tray>.

步驟 4:依場景挑選參考素材

最後,告訴模型每一個場景要用哪一個素材子集,還有場景結束時該長什麼樣:

Scene 1 | Truing the wheel in the workshop
Use: <Mechanic>, <Truing Stand>, <Workshop>, and the spoke-tightening motion from @Video 1.
Event: <Mechanic> spins the wheel in <Truing Stand> and tightens one spoke.
End state: <Mechanic> stands on the inner side of the bench. <Truing Stand> stays at the left of the frame.

Scene 2 | Handoff in the courtyard
Use: <Apprentice>, <Parts Tray>, and <Courtyard>.
Event: <Apprentice> carries <Parts Tray> to the courtyard bench and sets it down.
End state: <Parts Tray> rests on the bench. No other character enters the courtyard.

場景層級的挑選,就是把一堆參考素材變成一份鏡頭表的關鍵。你在某個場景裡沒有點名的每一份素材,模型都知道要把它留在外面。

音訊和文字的特殊語法

提示詞完全可以用自然語言寫,但當你需要把音樂、音效、對白和字幕分得清清楚楚時,Seedance 2.5 認得四種括號:

內容語法範例
音樂()(Soft, rhythmic piano music plays in the background)
音效<><A bell rings in the distance>
對白{}{Hello, welcome back.}
字幕【】【Chapter One: Departure】

把對白的語言再強調一次

當對白不是中文時,在台詞前面把語言講出來——例如 The girl says softly in Japanese: {もう大丈夫です}。如果寫出來的文字是英文、但你要模型用另一種語言唸出來,或是你需要某個特定的地區口音,就用這條公式:

對白語言 + 地區腔調或口音 + 演繹風格 + 說話者 + {dialogue}

Dialogue language: American English. The girl says in natural, conversational American English: {I thought you weren't coming.}

Dialogue language: authentic Los Angeles English. The young man says in natural Los Angeles vernacular: {No way, you actually made it.}

不要把參考素材已經定義好的東西再講一次

一個常見的寫太多的習慣:上傳了一段動態參考影片,然後又用文字把同一個動作再描述一遍。當一段參考影片已經準確定義了動作、運鏡和順序時,只要講出要繼承哪些屬性就好。重複描述那個動作可能會跟參考素材本身衝突,模型就得去調和同一句指令的兩個版本。

例外是預演影片。預演主要提供的是動態和空間結構——路徑、走位、運鏡、剪接——所以提示詞還是得定義預期的主體、場景、動作和視覺風格。時間點跟影片繼承;世界由圖片和文字提供。如果預演參考素材是你的主要流程,Seedance 2.5 分鏡與預演指南對粗預演和細預演談得更深。

送出之前把清單跑一遍

按下生成之前,拿這些問題讀一遍你的提示詞:

  • 提示詞有沒有清楚講出主體和主要動作或事件?
  • 每一份參考輸入有沒有講出要用什麼、不要用什麼?
  • 每一個不同的角色、產品和道具,有沒有被點名並綁到一份參考素材上?
  • 參考素材是依場景挑選的,還是被要求全部一次出現?
  • 角色數量、服裝、道具歸屬和空間關係,有沒有維持一致?
  • 抽象的情緒和攝影術語,有沒有配上看得見或聽得見的線索?

注意

參考素材提示詞提高的是「結果忠實」的機率;它不保證忠實。字幕、公式、招牌、產品規格,或必須精確到影格的時間點,都要把準備好的參考輸入、生成和後期製作結合起來。

在 OmniArt 開始

打開 OmniArt 的影片創作工作台,選 Seedance 2.5,照這個順序寫出你的第一句多參考素材提示詞:

  1. 只上傳那些帶著別的素材沒有的資訊的素材。
  2. 每一個主體寫一行對應——名字、素材,還有要從它身上拿什麼。
  3. 只要有可能滲出背景、人或構圖的那幾行,都加上一句排除項。
  4. 把對應關係分組放在 [Characters][Props][Scenes][Motion and Audio] 底下。
  5. 替任何跨多個場景的角色寫一份主體檔案。
  6. 逐場景列出它的素材子集、它的事件,還有它的結束狀態。
  7. 用同一句提示詞生成兩次,然後在兩次結果不一致的地方去修對應關係——不是修形容詞。

想知道這次釋出改了什麼、新模式該放在哪裡,可以讀 Seedance 2.5 上線了什麼。想把同一套參考素材紀律帶到更長的作品裡,可以接著看怎麼替 30 秒的 Seedance 2.5 影片寫提示詞;想要能貫穿一整個專案的識別控制,可以看我們的AI 影片角色一致性指南

準備好開始創作了嗎?

用 AI 生成精彩內容

免費開始