Seedance 2.5 reference prompting: give each material a role
Seedance 2.5 reference prompting, step by step: assign every image, video, and audio clip a role, bind subjects to materials, and select references per scene.

Seedance 2.5 reference prompting works differently from the text-to-video prompting most creators learned first. The model accepts up to 50 reference inputs in one generation — images, videos, and audio combined — and once you are past two or three of them, the hard part is no longer describing a scene. It is telling the model which material means what, and which ones belong in each moment of the video.
That shift is worth naming plainly. With one reference image, the model can guess your intent. With twelve, guessing produces blended faces, swapped clothing, duplicated props, and backgrounds bleeding in from images you only uploaded for a face. The fix is not a longer scene description. It is a short, explicit block at the top of your prompt that assigns every material a job and a set of exclusions.
Seedance 2.5 is available in the OmniArt video creation workspace, so you can follow along with your own references as you read. This guide covers the core prompt formula, the material limits, the role-definition pattern, and the four-step workflow for prompts that carry many subjects at once.
Start from the core prompt formula
Before references enter the picture, Seedance 2.5 responds to a simple, flexible formula. Every element after the first two is optional — include what changes the shot, drop what does not.
Subject + action or event + scene and environment + visual style + camera movement or cut + audio
Written out as a template:
<Subject> performs <primary action or event> in <scene and environment>.
The visuals feature <visual style>.
Use <shot size, camera angle, camera movement, or cuts>.
Audio includes <dialogue, ambience, sound effects, or music>.
And filled in for a real shot:
A bicycle mechanic trues a wheel at a workbench in a narrow street-front workshop, then lifts the wheel onto a wall hook.
Late afternoon light comes through the roll-up door. Metal surfaces are worn and slightly oily, and the floor is swept.
Begin with a medium shot of the truing stand, push in slowly toward the spoke nipples, then cut to a frontal view of the wall.
Retain the ping of spoke tension, the click of the truing gauge, and quiet street ambience.
Generation parameters such as resolution and duration do not belong in the prompt — set those on the generation page. Everything in the prompt should be content the model has to interpret.
Know your material budget before you write
Seedance 2.5 combines up to 50 reference inputs across three types. Each type has a hard input limit and a separate recommended range. The recommended range is not a cap — it is the zone where generations stay stable.
| Material type | Input limit | Recommended range |
|---|---|---|
| Images | Up to 30 images, each no larger than 4K | 1–8 distinct subjects across subject-reference images |
| Videos | Up to 10 videos, combined duration no more than 30 seconds | 1–5 distinct subjects, 5–10 seconds per subject video |
| Audio | Up to 10 clips, combined duration no more than 30 seconds | Only dialogue, voice, ambience, or music relevant to the task |
| Video editing | One source video plus reference images | Source video under 20 seconds, 1–5 reference images |
You can push past the recommendations — 9–12 subjects in subject images, 6–10 subjects in subject audio or video, 6–8 reference images for an edit. Stability tends to decrease as the count grows, so treat anything above the recommended range as a deliberate experiment rather than a default.
One packing rule matters more than people expect: if more than five subjects also need multiple views, put each view in its own image. Independent view images are usually more stable than several views combined into one collage sheet. A four-panel contact sheet reads to the model as one picture with four things in it, which is exactly the ambiguity you are trying to remove.
Tip
Trim before you upload. Ten tight, single-purpose materials outperform twenty overlapping ones, because every extra material is another thing the model has to disambiguate.
Define each material's role — and its exclusions
This is the technique the rest of the guide builds on. After uploading, state exactly what each material contributes, and state what must not be carried over. Material mappings belong in the prompt text; do not rely on labels burned into the images, and do not leave the model to infer which person or prop a file represents.
The pattern:
@Image 1 defines <subject>'s <appearance, clothing, structure, or material>.
@Video 1 defines <motion, camera movement, or pacing>.
@Audio 1 defines <character or sound type>'s <voice, dialogue, ambience, or music>.
<Subject> completes <primary action or event> in <scene>.
The visuals feature <visual style>, with <camera treatment>.
Filled in, with exclusions doing real work:
@Image 1 defines the mechanic's facial features, hairstyle, and navy work apron. Do not use the image background.
@Image 2 defines the workshop's bench layout, tool wall, and roll-up door light. Do not use the people in the image.
@Video 1 defines the pacing of spinning the wheel, pausing, and tightening a spoke. Do not use the person's identity, clothing, or scene from the video.
The mechanic trues a wheel at the workbench, then lifts it onto the wall hook.
Begin with a medium shot of the truing stand and push in slowly toward the spoke nipples. Retain spoke pings and quiet street ambience.
The exclusion lines are not padding. A portrait reference carries a background, a lighting direction, and often other people. Without "do not use the image background," those attributes are fair game for the model, and they will show up somewhere.
Bind multiple views to one object
When several images show the same person or product from different angles, say so explicitly, and say how many of the object should exist in the output:
@Image 1 defines the front view of the same steel touring frame.
@Image 2 defines the left-side structure of the same steel touring frame.
@Image 3 defines the right-side structure of the same steel touring frame.
@Image 4 defines the rear structure of the same steel touring frame.
All four images define one steel touring frame. The output must contain only one frame throughout.
Without that closing sentence, four images of one object is a plausible description of four objects. The count assertion is the cheapest insurance in the whole prompt.
The four-step multi-reference workflow
Once you are past a handful of materials, work through these four steps in order. The goal is to help the model pick the right materials for the current scene — not to make everything appear at once.
Step 1: Name and map each subject individually
Bind every person, product, and prop to its own material, one line each:
<Character A> corresponds to @Image 1. Use only the appearance, hairstyle, and clothing.
<Character B> corresponds to @Image 2. Use only the appearance, hairstyle, and clothing.
<Prop A> corresponds to @Image 3. Use only the structure, material, and color.
<Scene A> references @Image 4. Use only the spatial layout, architecture, and lighting. Do not use the people in the image.
The anti-pattern to avoid: never write "@Images 1 through 4 define four characters respectively." That sentence tells the model there are four characters and four images, but never says which is which — so it is free to pair them however it likes.
Step 2: Group materials by type
Once each mapping exists, organize them under headers so relationships are readable at a glance:
[Characters]
<Mechanic> corresponds to @Image 1. Use only the appearance, hairstyle, and clothing.
<Apprentice> corresponds to @Image 2. Use only the appearance, hairstyle, and clothing.
Do not interchange the two characters' appearances, clothing, actions, positions, or dialogue.
[Props]
<Truing Stand> corresponds to @Image 3 and belongs only to <Mechanic>.
<Parts Tray> corresponds to @Image 4 and belongs only to <Apprentice>.
[Scenes]
<Workshop> references @Image 5. Use only the space, materials, and lighting.
<Courtyard> references @Image 6. Use only the space, materials, and lighting.
[Motion and Audio]
@Video 1 defines the motion of <Mechanic> tightening a spoke. Do not use the person or scene from the video.
@Audio 1 defines <Apprentice>'s voice and specified dialogue.
Prop ownership lines matter as much as character lines. "Belongs only to" is what stops a tool from teleporting between two people between cuts.
Step 3: Write a profile for important subjects
When one character draws on several materials across multiple scenes, collect them in one block instead of leaving them scattered:
[Subject Profile: Mechanic]
Appearance and clothing: @Image 1.
Fixed prop: <Truing Stand> from @Image 3.
Locations: <Workshop> and <Courtyard>.
Motion references: the spoke-tightening motion from @Video 1 and the wheel-lifting motion from @Video 2.
Do not use: the apprentice's clothing. Do not give this character <Parts Tray>.
Step 4: Select references by scene
Finally, tell the model which subset of materials each scene uses, and what the scene should look like when it ends:
Scene 1 | Truing the wheel in the workshop
Use: <Mechanic>, <Truing Stand>, <Workshop>, and the spoke-tightening motion from @Video 1.
Event: <Mechanic> spins the wheel in <Truing Stand> and tightens one spoke.
End state: <Mechanic> stands on the inner side of the bench. <Truing Stand> stays at the left of the frame.
Scene 2 | Handoff in the courtyard
Use: <Apprentice>, <Parts Tray>, and <Courtyard>.
Event: <Apprentice> carries <Parts Tray> to the courtyard bench and sets it down.
End state: <Parts Tray> rests on the bench. No other character enters the courtyard.
Scene-level selection is what converts a pile of references into a shot list. Every material you do not name in a scene is a material the model knows to leave out.
Special syntax for audio and text
Prompts can be written entirely in natural language, but when you need music, sound effects, dialogue, and subtitles kept distinct, Seedance 2.5 recognizes four bracket types:
| Content | Syntax | Example |
|---|---|---|
| Music | () | (Soft, rhythmic piano music plays in the background) |
| Sound effects | <> | <A bell rings in the distance> |
| Dialogue | {} | {Hello, welcome back.} |
| Subtitles | 【】 | 【Chapter One: Departure】 |
Reinforce the dialogue language
When dialogue is not in Chinese, name the language before the line — for example, The girl says softly in Japanese: {もう大丈夫です}. If the written text is English but the model speaks it in another language, or you need a particular regional variety, use this formula:
Dialogue language + regional variety or accent + delivery style + speaker + {dialogue}
Dialogue language: American English. The girl says in natural, conversational American English: {I thought you weren't coming.}
Dialogue language: authentic Los Angeles English. The young man says in natural Los Angeles vernacular: {No way, you actually made it.}
Don't restate what a reference already defines
A common over-writing habit: uploading a motion reference video and then describing the same motion again in prose. When a reference video already defines the movement, camera work, and sequence accurately, state only which attributes to inherit. Repeating the motion description can conflict with the reference itself, and the model has to reconcile two versions of one instruction.
The exception is a blockout video. A blockout mainly supplies motion and spatial structure — paths, blocking, camera movement, cuts — so the prompt still has to define the intended subjects, scene, action, and visual style. Inherit the timing from the video; supply the world from images and text. If blockout references are your main workflow, the Seedance 2.5 storyboards and blockouts guide goes deeper on coarse versus fine blockouts.
Run the checklist before you submit
Before you press generate, read your prompt against these questions:
- Does the prompt clearly state the subject and the primary action or event?
- Does every reference input state what to use and what not to use?
- Is every distinct character, product, and prop named and bound to a reference?
- Are references selected by scene rather than required to appear all at once?
- Do the number of characters, clothing, prop ownership, and spatial relationships stay consistent?
- Are abstract emotions and cinematography terms paired with visible or audible cues?
Warning
Reference prompting raises the probability of a faithful result; it does not guarantee one. For subtitles, formulas, signage, product specifications, or frame-level timing that must be exact, combine prepared reference inputs, generation, and post-production.
Getting started on OmniArt
Open the OmniArt video creation workspace, pick Seedance 2.5, and build your first multi-reference prompt in this order:
- Upload only the materials that carry information no other material carries.
- Write one mapping line per subject — name, material, and what to use from it.
- Add an exclusion to every line where a background, person, or composition could leak.
- Group the mappings under
[Characters],[Props],[Scenes], and[Motion and Audio]. - Write a subject profile for any character that spans multiple scenes.
- List each scene with its material subset, its event, and its end state.
- Generate twice with the same prompt, then fix mappings — not adjectives — wherever the two runs disagree.
For what changed in this release and where the new modes fit, read what shipped in Seedance 2.5. To take the same reference discipline into longer pieces, continue with prompting 30-second Seedance 2.5 videos, and for identity control that carries across a whole project, see our guide to consistent characters in AI video.
Ready to create?
Start generating amazing content with AI