Seedance 2.5 Prompt Guide for Professionals
Seedance 2.5 takes direction unusually well. It accepts text, up to 50 reference files, and briefs for clips up to 30 seconds, and it follows structure closely enough that how you write the prompt decides how much of your shot survives generation.
This guide covers how to write for it. You can run Seedance 2.5 on mitte directly in the browser.
One rule before anything else: generation settings (duration, resolution, aspect ratio, audio on or off) are controls on the tool bar in mitte. Keep them out of the prompt. The prompt is for the shot.
Start with the formula
Every Seedance prompt is some subset of six elements:
Subject + action + scene + visual style + camera + audio.
Subject and action are the foundation: who or what is doing what. Scene sets location, time, and weather. Style covers lighting, color, and texture. Camera covers shot size, angle, movement, and cuts. Audio covers dialogue, ambience, effects, and music. Drop whatever you don’t need.
A ceramic artist finishes a pale blue cup in a studio at dawn, lifts it from the wheel, and places it in the center of a wooden shelf. Soft morning light enters through the window. The wet clay has a delicate sheen. Begin with a medium shot of the wheel-throwing process, slowly push in toward the cup’s surface texture, then cut to a frontal view of the shelf. Retain the low hum of the pottery wheel, the friction of clay, and subtle indoor ambience.
Notice the shape: one continuous event, described in the order a viewer would see it.
Give every reference a job
Seedance 2.5 accepts up to 30 images, 10 videos, and 10 audio clips in a single run (50 files total, with combined video and audio under 30 seconds). Upload them in the prompt drop zone and tag them as @Image1, @Video1, @Audio1.
Two rules make references behave:
- State what each file contributes. Appearance, motion, pacing, voice, spatial layout. Be exact.
- State what to ignore. Backgrounds, bystanders, and compositions leak into the output unless you exclude them.
@Image 1 defines the ceramic artist's facial features, hairstyle, and dark green apron. Do not use the image background.
@Image 2 defines the wooden workbench, window placement, and morning light of the pottery studio. Do not use the people in the image.
@Video 1 defines the pacing of throwing clay with both hands, lifting the cup, and placing it down. Do not use the person's identity, clothing, or scene from the video.
The ceramic artist finishes a pale blue cup in the pottery studio at dawn, lifts it from the wheel, and places it in the center of a wooden shelf.
The mapping has to live in the prompt. Text labels inside the images do nothing, and the model won’t guess which face belongs to which character.
If several images show different views of one subject, say so, or you’ll get four lamps:
@Image 1 defines the front view of the same folding desk lamp.
@Image 2 defines the left-side structure of the same folding desk lamp.
All images define one folding desk lamp. The output must contain only one lamp throughout.
Separate view images beat a collage. When a subject needs multiple angles, upload them as individual files.
The audio and text syntax
Plain language works for sound, but four bracket types remove ambiguity:
| Content | Syntax | Example |
|---|---|---|
| Music |
() |
(Soft, rhythmic piano plays in the background) |
| Sound effects |
<> |
\<A bell rings in the distance> |
| Dialogue |
{} |
{Hello, welcome back.} |
| Subtitles |
【】 |
【Chapter One: Departure】 |
If dialogue comes out in the wrong language or accent, reinforce it before the line:
Dialogue language: American English. The girl says in natural, conversational American English: {I thought you weren’t coming.}
Directing a cast
With many references, the job shifts from describing files to defining relationships: which character owns which prop, which scene uses which materials. Bind each subject individually.
[Characters]
<Conservator> corresponds to @Image 1. Use only the appearance, hairstyle, and clothing.
<Registrar> corresponds to @Image 2. Use only the appearance, hairstyle, and clothing.
[Props]
<Sample Case> corresponds to @Image 3 and belongs only to <Conservator>.
[Scenes]
<Conservation Lab> references @Image 4. Use only the space, materials, and lighting.
Never write “@Images 1 through 4 define four characters respectively.” That sentence binds nothing, and the model will shuffle faces.
For a multi-scene piece, select references per scene instead of demanding everything at once:
Scene 1 | Inspection in the Conservation Lab
Use: <Conservator>, <Sample Case>, <Conservation Lab>.
Event: <Conservator> opens <Sample Case> at the workbench and inspects the sample inside.
End state: <Conservator> remains on the inner side of the workbench, case beside the right hand.
The goal of multi-reference prompting is picking the right materials for each moment. Files that belong to another scene should sit out until that scene arrives.
Long videos: one change per stage
A 30-second clip is several events, and Seedance handles them best as consecutive stages. Give each stage exactly one primary state change and end it with something directly visible.
[Stage 1]
Initial state: <Florist> stands behind the workbench. Loose stems, scissors, and wrapping paper on the tabletop.
Primary event: <Florist> arranges the stems and trims them to length.
End state: <Florist> holds the bouquet in the left hand, scissors back on the right side of the workbench.
[Stage 2]
Continue from the previous stage: same identities and clothing, <Florist> still holds the bouquet.
Primary event: <Store Assistant> unfolds the wrapping paper. <Florist> places the bouquet inside and ties it with a green ribbon.
End state: the wrapped bouquet lies flat in the center of the workbench, ribbon bow facing the camera.
End states are the load-bearing part. They tell the model what must be true on screen before the next stage begins, which is what keeps props from teleporting.
Timestamps work too (0-5 seconds, 5-10 seconds), but treat them as time budgets, never as frame-accurate edit points. Use an exact time only for a single critical beat: “At 5 seconds, the camera whip-pans rapidly to the left.” Don’t use them to demand frequencies like three actions in one second.
First frames, last frames, keyframes
In mitte, flip the mode toggle to frame mode for a plain first-frame or first-and-last-frame run. In omni mode you can do the same inline, with room for extra references:
@Image 1 is the first frame. It defines the opening composition, subject position, pose, and camera direction.
@Image 2 is the last frame. It defines the ending composition.
@Image 3 defines <Perfumer>'s face, hairstyle, and dark green apron. Do not change the compositions defined by @Image 1 and @Image 2.
Describe each anchor image separately (never “@Images 1 and 2 are the first and last frames”), and give the first and last images the same aspect ratio or the final frame stretches.
The same pattern scales to a sequence: “Use @Image 1 through @Image 4 as keyframes in this order,” then describe the visible state each keyframe represents. Keyframes control stage order and key states; they don’t reproduce every pixel.
A storyboard grid works as a reference too: tell the model the reading order and warn it off the line-art style. And a gray blockout video can carry paths, blocking, and camera moves while separate images define what everything actually looks like. Map each geometric shape to its final subject (“The tall cylinder in @Video 1 corresponds to <Guide>”).
Editing and extending existing videos
Upload a source video as a reference and Seedance 2.5 will edit it. The structure that keeps edits surgical: declare the master, scope the change, list what survives.
[Edit Goal]
Edit @Video 1. Replace only the yellow folding desk lamp with the white folding desk lamp in @Image 1.
[Source Video Role]
@Video 1 is the sole editing master. It defines the desk, hand movements, camera movement, occlusion, and event order.
[Edit Scope]
Keep exactly one white folding desk lamp throughout. Do not modify the books, desk, hands, or background.
[Timeline Inheritance]
The white lamp inherits every appearance, rotation, hand occlusion, and exit of the original, including timing, path, and speed changes.
Edits preserve the source video’s aspect ratio and roughly its duration (expect up to about 0.3 seconds of drift). Audio edits work the same way: name the sound category, the change, and everything that must stay.
Extensions hinge on the boundary frame. Extending forward, describe the source’s last frame as the continuous starting state, then the new action. Extending backward, describe what happens before, then define the source’s first frame as the explicit end state. “Then connect to the source video” on its own lets later characters wander in early.
Write emotion you can see
“Tense” and “warm” set a direction, and the model fills the performance however it likes. For control, trade adjectives for 2 to 4 observable cues: eye movement, brow tension, breathing, hand position.
After confirming that the curtain call is over, the actor exhales softly. The shoulders gradually relax, a restrained smile appears, and the eyes slowly well with tears, but the actor never turns to leave.
That’s a whole emotional arc without a single emotion word.
Camera language
Standard terms go straight in: wide shot, close-up, push in, orbit, low angle, first-person view, dolly zoom, whip pan, handheld. For anything niche, keep the term and translate it into a visible change:
Rack focus: shift focus smoothly from the leaves in the foreground to the person in the background. The leaves gradually blur while the person’s face changes from soft to sharp.
When several subjects share the frame, say which one the camera follows and where the move starts and ends.
Before you hit generate
- Subject and primary action stated?
- Every reference told what to use and what to ignore?
- Every character, product, and prop named and bound to a file?
- One primary change and a visible end state per stage?
- For edits: sole master declared, scope and preserved content listed?
- Emotions and niche camera terms paired with visible cues?
- First and last frames on the same aspect ratio?
Open Seedance 2.5 on mitte and direct something.