All guides
Seedance 2.5 Guide

Seedance 2.5 Prompt Guide

Learn how to turn a rough idea into a controllable Seedance 2.5 prompt—then extend the same method to references, 30-second scenes, video edits, transitions, audio, and performance direction.

The short version

Subject + action + setting + look + camera + sound

That formula is a starting point, not a form you must fill in. A clear subject and event matter most. Add the remaining layers only when they change what should appear on screen or be heard in the finished video.

On this page
  1. 01Core prompt formula
  2. 02Reference materials
  3. 03Audio and dialogue
  4. 04Long videos and timing
  5. 05Editing and extension
  6. 06Advanced control
  7. 07Final checklist
  8. 08FAQ
Section 01

Start with a simple Seedance 2.5 prompt formula

The strongest prompts read like a short production brief. They tell Seedance 2.5 what changes over time, not just what a still frame looks like. Begin with the event; then add the visual and audio decisions that help stage it.

Subject and event

Who or what is present, and what changes during the shot.

Scene

Place, time, weather, background state, and useful spatial relationships.

Visual treatment

Lighting, color, texture, materials, realism, and overall mood.

Camera

Framing, angle, movement, focus, cuts, and the subject the camera follows.

Audio

Dialogue, voice, ambience, effects, and music that belong to the action.

You do not need every ingredient in every Seedance 2.5 prompt. If a reference already gives you the look or the camera move, say what to inherit and spend the prompt on the action that is still undefined. Resolution, aspect ratio, and duration belong in the generation controls rather than the prose prompt.

Text-to-video example
A compact electric motorcycle follows a rain-darkened mountain road before sunrise. The rider leans into one broad curve as mist moves across the valley below.

Use cool blue ambient light with a thin warm line on the horizon. Keep the motorcycle geometry stable and the wet asphalt reflective.

Begin with a low rear three-quarter tracking shot, move alongside the rider through the curve, then widen to reveal the valley.

Audio: restrained motor whine, tire spray, light wind, and distant birds. No music.

Put positive direction before negative constraints. Describe the desired shot clearly, then add a short “keep” or “do not” line for the failure modes that matter most.

Section 02

Give every reference one clear job

Uploading more material does not automatically create more control. Seedance 2.5 works best when the prompt explains what each image, video, or audio clip contributes—and what should be ignored.

ByteDance’s official Dreamina guide describes support for up to 50 reference assets on compatible Seedance 2.5 surfaces. The documented breakdown is up to 30 images, 10 videos with no more than 30 seconds combined, and 10 audio clips with no more than 30 seconds combined. The recommended ranges below are guidance for generation stability, not lower capability limits.

Material typeOfficial input limitRecommended range
ImagesUp to 30 images, each no larger than 4KPrefer 1–8 distinct subjects across subject-reference images
VideosUp to 10 videos, with a combined duration of 30 seconds or lessPrefer 1–5 distinct subjects and 5–10 seconds per subject-reference video
AudioUp to 10 audio clips, with a combined duration of 30 seconds or lessKeep only dialogue, voice traits, ambience, or music that directly serves the task
Video editingOne source video may be used together with reference imagesPrefer a source video under 20 seconds and 1–5 reference images

You can go beyond those preferred ranges—for example, 9–12 subjects in subject images, 6–10 subjects represented through audio or video, or 6–8 reference images in a video-editing job. The tradeoff is stability: as the number of materials and relationships grows, identity, motion, and scene assignments become harder to preserve consistently.

When more than five subjects also need multiple views, upload each view as its own image. Separate front, side, and rear views are generally more stable than a collage that compresses several views into one image.

Provider and product limits can differ. Treat the current Inspix upload controls as the source of truth for the workflow you are using; the table above summarizes ByteDance’s official guide.

Map references before describing the shot

Bind each important person, product, prop, location, and sound separately. Write the mapping in the prompt itself; do not rely only on labels embedded inside an image or expect the model to infer which person, prop, or scene a file represents. Name the subjects, assign each source, and add exclusions when a source contains unwanted people, backgrounds, or composition cues.

If several images show one product from different sides, say that they describe the same single product. If a motion clip is useful only for movement, explicitly reject its person, clothing, and background. This is how you prevent a helpful reference from quietly changing the rest of the scene.

Start with a reusable reference-role template

Reference role template
@Image 1 defines <Subject A>'s <appearance, clothing, structure, or material>. Do not inherit <unwanted background, people, or composition>.
@Video 1 defines <motion, camera movement, or pacing>. Do not inherit <identity, wardrobe, product design, or setting>.
@Audio 1 defines <speaker or sound category>'s <voice, dialogue, ambience, effects, or music>.

<Subject A> completes <primary action or event> in <scene>.
Use <visual treatment> with <shot size, camera angle, movement, or cuts>.

Then turn those roles into a complete shot

Multimodal reference example
[REFERENCE ROLES]
@Image 1 defines the rider's face, helmet, and charcoal riding suit. Do not inherit its background.
@Image 2 defines the motorcycle's frame, headlight shape, and red side panels. Treat every view as the same single motorcycle.
@Video 1 defines the rider's lean, road speed, and the camera's side-tracking motion. Do not inherit the person, vehicle design, or location from the video.
@Audio 1 defines the quiet electric motor tone and tire spray.

[SHOT]
The rider follows a wet mountain road at blue hour and leans through one wide curve. Keep the rider's identity, clothing, motorcycle structure, road direction, and weather consistent. End on a wide view with the rider continuing toward the bright horizon.

State explicitly when several images show one subject

Multiple angles should reinforce one identity or object, not multiply it. Spell out what each view contributes and confirm that all views belong to the same single subject.

Multi-view reference example
@Image 1 defines the front view of the same portable projector.
@Image 2 defines the left-side controls of the same portable projector.
@Image 3 defines the right-side vents of the same portable projector.
@Image 4 defines the rear ports of the same portable projector.

All four images describe one projector. Keep exactly one projector in the video and preserve the same housing, proportions, controls, vents, and ports from every angle.

If a reference video already defines the motion, camera path, and event order accurately, describe only the attributes to inherit. Repeating every action in prose can conflict with the reference. A coarse blockout mainly supplies motion and spatial structure, so the prompt must still define the intended subjects, setting, action, and visual style.

Reference tokens may be displayed as @Image 1, @Image1, or a visual mention chip depending on the product interface. Keep the exact token inserted by the uploader; the important part is the explicit role that follows it.

A reliable workflow for many references

  1. STEP 1

    Map each subject

    Give every recurring character, product, prop, and location a stable name and a specific reference.

  2. STEP 2

    Group by role

    Keep character identity, object design, scene, motion, and audio references logically separated.

  3. STEP 3

    Build a subject profile

    For an important recurring subject, collect its appearance, clothing, fixed props, locations, and motion sources in one block.

  4. STEP 4

    Select per scene

    For each scene, list only the subjects and references needed there, plus the event and visible end state.

Section 03

Write audio and dialogue so each layer is unambiguous

Natural language is enough for most Seedance 2.5 prompts. When a scene contains several kinds of audio or visible text, lightweight syntax can separate them and reduce confusion.

Music(slow analog synth under the scene)
Sound effect<a metal latch clicks shut>
Dialogue{We have one minute left.}
Subtitle【Field Test — Day 03】

For non-Chinese dialogue, state the language before the line. This is especially useful when English text is spoken in Chinese by default or when the performance needs a specific regional variety. Add the accent or regional variety only when it matters, followed by delivery style and the speaker.

Dialogue language formula

Language + regional variety or accent + delivery style + speaker + {Dialogue}

Dialogue language example
Dialogue language: British English.
Regional variety or accent: contemporary London English.
Delivery: quiet, hurried, and natural.
Speaker: the station clerk.
Dialogue: {The last train leaves in two minutes.}

Do not ask every audio layer to dominate. Decide what is foreground sound, what sits underneath, and what must remain unchanged when editing. A short hierarchy produces a cleaner result than a long list of equally loud cues.

Section 04

Structure 30-second prompts as stages, not a dense paragraph

Longer clips fail when too many actions compete at once. Break the sequence into consecutive stages, give each stage one main change, and describe what the viewer should see when that stage ends.

An end state is more useful than “then continue naturally.” It tells Seedance 2.5 where every important person and prop should be before the next event starts. End states are especially valuable for handoffs, assembly, entrances, exits, and scenes where object count must remain stable.

Multi-stage prompt example
[GOAL]
Create a 24-second product film showing a craftsperson assembling a portable desk lamp.

[STAGE 1 — PARTS]
Initial state: the lamp components are arranged separately on a clean wooden table.
Primary event: the craftsperson connects the base, arm, and light head in that order.
End state: the assembled lamp stands upright in the center; no loose parts remain in the hands.

[STAGE 2 — ADJUSTMENT]
Continue with the same person, table layout, and lamp structure.
Primary event: the arm extends and the light head rotates toward an open sketchbook.
End state: the beam lands on the center of the page and both hands leave the lamp.

[STAGE 3 — HERO SHOT]
The room becomes slightly darker while the lamp remains on. The camera makes a restrained push-in.
Final state: the lamp is centered, fully assembled, and unchanged in shape.

Maintain one lamp, stable materials, consistent hand count, continuous screen direction, and quiet workshop ambience throughout.

Use timestamps only when timing is part of the brief

Time range

Allocate a section of the runtime: “0–5s: establish the empty room.”

Exact moment

Reserve one beat: “At 8 seconds, the doors open.”

Relative timing

Link events: “Two seconds after the alarm, the lights turn red.”

Keep time ranges consecutive and non-overlapping. Treat them as pacing budgets, not frame-accurate edit points. Too little content leaves the model free to improvise; too much content encourages rushed cuts or skipped actions. For ordinary narrative work, stages should remain the default.

Section 05

Separate generation, editing, and extension logic

A new video prompt describes what should be created. An editing prompt must protect a source video. An extension prompt must first connect to the source boundary. Treating these as different jobs makes the instruction much easier to follow.

For video editing: master, target, scope, preserve

Declare one source clip as the editing master. It controls the characters, scene, action, composition, camera, timing, and audio unless you explicitly change one of them. Then identify the target, narrow the time or region, and list the elements that must survive the edit.

Video editing template
[EDIT GOAL]
Edit @Video 1. Between 4 and 7 seconds, change only the cool blue wall light to a warm amber light.

[SOURCE MASTER]
@Video 1 is the sole editing master for the character, room layout, action, framing, camera movement, event order, dialogue, and ambience.

[EDIT SCOPE]
Modify only the wall light and the surfaces it naturally illuminates. Let the skin and clothing respond subtly to the warmer light.

[PRESERVE]
Keep the character's identity, expression, position, motion, clothing, room structure, camera path, speech, and timing unchanged.

For extension: describe the boundary before the new event

A forward extension should begin from the source video’s final frame. A backward extension should finish at its first frame. Before adding a new action, describe the boundary state: subject pose, prop positions, scene geometry, camera direction, lighting, motion direction, and any continuous sound.

Forward extension template
[SOURCE]
@Video 1 is the video to extend forward.

[BOUNDARY]
Begin directly from the final frame of @Video 1. Preserve the subject's pose and direction, the bicycle's position, the street layout, the camera height, the afternoon light, and the current forward motion.

[NEW ACTION]
The cyclist exits the narrow street into an open plaza, slows beside a fountain, and looks toward the clock tower. The camera continues the existing tracking movement before easing into a wide reveal.

[CONTINUITY]
Keep the same cyclist, clothing, bicycle structure, travel direction, weather, and ambient city sound. Do not duplicate the cyclist or bicycle.

In input-led modes, aspect ratio or duration may be inherited from the source material. Use the generation interface for available settings and keep the prompt focused on content and continuity.

Section 06

Advanced control without overloading the prompt

Keyframes, storyboards, blockouts, transition clips, and performance notes solve different problems. State which control each input owns instead of asking every asset to define the whole result.

First and last frames

Identify the opening and ending images separately, keep their aspect ratios aligned, and describe one continuous event that travels between them. Additional images should supplement identity, props, or materials without replacing either anchor composition.

Multiple keyframes and storyboards

Declare the image order first, then explain the state represented by each frame. Separate keyframe images are usually clearer than a single crowded grid. A storyboard controls sequence and key states; it is not a promise to reproduce every panel pixel for pixel.

Coarse and fine blockouts

Use a coarse blockout for motion paths, blocking, entrances, camera movement, cuts, lighting changes, and audio rhythm. Use a clean fine blockout when structure is already solved and the job is to re-render materials, colors, characters, scene detail, or style.

Image-to-video assembly

When turning a group of images into one finished video, specify each image's role, the intended order, how much motion to add, editing rhythm, visual packaging, and sound. “Make these images into a video” leaves too many production decisions undefined.

Seamless transitions

Define the outgoing clip, incoming clip, trigger action, camera direction, visual transformation, arrival composition, and audio handoff. Occlusion, object morphs, focus changes, and continuous pushes work best when corresponding shapes and motion directions are clear.

Direct performance with visible behavior

Words such as “tense,” “warm,” or “relieved” describe a mood, but they do not fully stage an actor’s performance. Add two to four observable cues: gaze, brow tension, mouth movement, breathing, shoulders, hands, or speech delivery. Use event-led stages only when the emotion changes more than once.

Performance direction example
The performance moves from guarded concentration to quiet relief.

When the mechanic hears the repaired engine start, the hands stop above the open hood and the eyes shift toward the dashboard.

After the engine settles into a steady idle, the shoulders loosen, the breath releases slowly, and a small restrained smile appears.

Keep the acting natural and understated. Preserve the character's position, clothing, and hand placement.
Section 07

Seedance 2.5 prompt checklist

Before generating, read the prompt once as a production handoff. If a collaborator could not tell what changes, what stays fixed, or which reference controls which detail, the model will face the same ambiguity.

  • The main subject and event are stated in the opening lines.
  • Every reference has one named role and useful exclusions.
  • Different views are identified as the same person or product.
  • Long scenes are divided into stages with visible end states.
  • Time ranges are consecutive and contain a realistic amount of action.
  • Camera direction names the subject being followed or revealed.
  • Dialogue includes its language, speaker, and delivery when needed.
  • Edits define one master video, a narrow scope, and preserved content.
  • Extensions describe the boundary frame before introducing new events.
  • Constraints protect specific risks instead of becoming a generic negative list.
Section 08

Seedance 2.5 prompt FAQ

Quick answers to the questions that usually appear between a first draft and a controllable generation.

What is the best structure for a Seedance 2.5 prompt?+

Start with the subject and the main action. Add the setting, visual treatment, camera direction, and sound only when they matter. For a longer scene, replace one dense paragraph with consecutive stages and a visible end state for each stage.

Should every reference be described in the prompt?+

Yes. Give each reference one clear job, such as identity, product structure, motion, camera movement, ambience, or music. Also state what should not be inherited when a reference contains an unwanted person, setting, or composition.

Do Seedance 2.5 prompts need timestamps?+

Not always. Stages are easier to maintain for most narratives. Use time ranges when pacing must be allocated, an exact timestamp for one critical beat, or relative timing when one action must happen after another.

How should I prompt a Seedance 2.5 video edit?+

Declare the source video as the editing master, identify the precise object, region, time range, or audio layer to change, then list what must remain untouched. Narrow edit scopes usually produce more predictable results.

What is the difference between a coarse and fine blockout?+

A coarse blockout mainly communicates motion, blocking, camera paths, and timing. A fine blockout already carries reliable structure, so the prompt can focus on re-rendering materials, color, characters, scene detail, and style.