Start with a simple Seedance 2.5 prompt formula
The strongest prompts read like a short production brief. They tell Seedance 2.5 what changes over time, not just what a still frame looks like. Begin with the event; then add the visual and audio decisions that help stage it.
Subject and event
Who or what is present, and what changes during the shot.
Scene
Place, time, weather, background state, and useful spatial relationships.
Visual treatment
Lighting, color, texture, materials, realism, and overall mood.
Camera
Framing, angle, movement, focus, cuts, and the subject the camera follows.
Audio
Dialogue, voice, ambience, effects, and music that belong to the action.
You do not need every ingredient in every Seedance 2.5 prompt. If a reference already gives you the look or the camera move, say what to inherit and spend the prompt on the action that is still undefined. Resolution, aspect ratio, and duration belong in the generation controls rather than the prose prompt.
A compact electric motorcycle follows a rain-darkened mountain road before sunrise. The rider leans into one broad curve as mist moves across the valley below. Use cool blue ambient light with a thin warm line on the horizon. Keep the motorcycle geometry stable and the wet asphalt reflective. Begin with a low rear three-quarter tracking shot, move alongside the rider through the curve, then widen to reveal the valley. Audio: restrained motor whine, tire spray, light wind, and distant birds. No music.
Put positive direction before negative constraints. Describe the desired shot clearly, then add a short “keep” or “do not” line for the failure modes that matter most.
Give every reference one clear job
Uploading more material does not automatically create more control. Seedance 2.5 works best when the prompt explains what each image, video, or audio clip contributes—and what should be ignored.
ByteDance’s official Dreamina guide describes support for up to 50 reference assets on compatible Seedance 2.5 surfaces. The documented breakdown is up to 30 images, 10 videos with no more than 30 seconds combined, and 10 audio clips with no more than 30 seconds combined. The recommended ranges below are guidance for generation stability, not lower capability limits.
| Material type | Official input limit | Recommended range |
|---|---|---|
| Images | Up to 30 images, each no larger than 4K | Prefer 1–8 distinct subjects across subject-reference images |
| Videos | Up to 10 videos, with a combined duration of 30 seconds or less | Prefer 1–5 distinct subjects and 5–10 seconds per subject-reference video |
| Audio | Up to 10 audio clips, with a combined duration of 30 seconds or less | Keep only dialogue, voice traits, ambience, or music that directly serves the task |
| Video editing | One source video may be used together with reference images | Prefer a source video under 20 seconds and 1–5 reference images |
You can go beyond those preferred ranges—for example, 9–12 subjects in subject images, 6–10 subjects represented through audio or video, or 6–8 reference images in a video-editing job. The tradeoff is stability: as the number of materials and relationships grows, identity, motion, and scene assignments become harder to preserve consistently.
When more than five subjects also need multiple views, upload each view as its own image. Separate front, side, and rear views are generally more stable than a collage that compresses several views into one image.
Provider and product limits can differ. Treat the current Inspix upload controls as the source of truth for the workflow you are using; the table above summarizes ByteDance’s official guide.
Map references before describing the shot
Bind each important person, product, prop, location, and sound separately. Write the mapping in the prompt itself; do not rely only on labels embedded inside an image or expect the model to infer which person, prop, or scene a file represents. Name the subjects, assign each source, and add exclusions when a source contains unwanted people, backgrounds, or composition cues.
If several images show one product from different sides, say that they describe the same single product. If a motion clip is useful only for movement, explicitly reject its person, clothing, and background. This is how you prevent a helpful reference from quietly changing the rest of the scene.
Start with a reusable reference-role template
@Image 1 defines <Subject A>'s <appearance, clothing, structure, or material>. Do not inherit <unwanted background, people, or composition>. @Video 1 defines <motion, camera movement, or pacing>. Do not inherit <identity, wardrobe, product design, or setting>. @Audio 1 defines <speaker or sound category>'s <voice, dialogue, ambience, effects, or music>. <Subject A> completes <primary action or event> in <scene>. Use <visual treatment> with <shot size, camera angle, movement, or cuts>.
Then turn those roles into a complete shot
[REFERENCE ROLES] @Image 1 defines the rider's face, helmet, and charcoal riding suit. Do not inherit its background. @Image 2 defines the motorcycle's frame, headlight shape, and red side panels. Treat every view as the same single motorcycle. @Video 1 defines the rider's lean, road speed, and the camera's side-tracking motion. Do not inherit the person, vehicle design, or location from the video. @Audio 1 defines the quiet electric motor tone and tire spray. [SHOT] The rider follows a wet mountain road at blue hour and leans through one wide curve. Keep the rider's identity, clothing, motorcycle structure, road direction, and weather consistent. End on a wide view with the rider continuing toward the bright horizon.
State explicitly when several images show one subject
Multiple angles should reinforce one identity or object, not multiply it. Spell out what each view contributes and confirm that all views belong to the same single subject.
@Image 1 defines the front view of the same portable projector. @Image 2 defines the left-side controls of the same portable projector. @Image 3 defines the right-side vents of the same portable projector. @Image 4 defines the rear ports of the same portable projector. All four images describe one projector. Keep exactly one projector in the video and preserve the same housing, proportions, controls, vents, and ports from every angle.
If a reference video already defines the motion, camera path, and event order accurately, describe only the attributes to inherit. Repeating every action in prose can conflict with the reference. A coarse blockout mainly supplies motion and spatial structure, so the prompt must still define the intended subjects, setting, action, and visual style.
Reference tokens may be displayed as @Image 1, @Image1, or a visual mention chip depending on the product interface. Keep the exact token inserted by the uploader; the important part is the explicit role that follows it.
A reliable workflow for many references
- STEP 1
Map each subject
Give every recurring character, product, prop, and location a stable name and a specific reference.
- STEP 2
Group by role
Keep character identity, object design, scene, motion, and audio references logically separated.
- STEP 3
Build a subject profile
For an important recurring subject, collect its appearance, clothing, fixed props, locations, and motion sources in one block.
- STEP 4
Select per scene
For each scene, list only the subjects and references needed there, plus the event and visible end state.
Write audio and dialogue so each layer is unambiguous
Natural language is enough for most Seedance 2.5 prompts. When a scene contains several kinds of audio or visible text, lightweight syntax can separate them and reduce confusion.
(slow analog synth under the scene)<a metal latch clicks shut>{We have one minute left.}【Field Test — Day 03】For non-Chinese dialogue, state the language before the line. This is especially useful when English text is spoken in Chinese by default or when the performance needs a specific regional variety. Add the accent or regional variety only when it matters, followed by delivery style and the speaker.
Language + regional variety or accent + delivery style + speaker + {Dialogue}
Dialogue language: British English.
Regional variety or accent: contemporary London English.
Delivery: quiet, hurried, and natural.
Speaker: the station clerk.
Dialogue: {The last train leaves in two minutes.}Do not ask every audio layer to dominate. Decide what is foreground sound, what sits underneath, and what must remain unchanged when editing. A short hierarchy produces a cleaner result than a long list of equally loud cues.
Structure 30-second prompts as stages, not a dense paragraph
Longer clips fail when too many actions compete at once. Break the sequence into consecutive stages, give each stage one main change, and describe what the viewer should see when that stage ends.
An end state is more useful than “then continue naturally.” It tells Seedance 2.5 where every important person and prop should be before the next event starts. End states are especially valuable for handoffs, assembly, entrances, exits, and scenes where object count must remain stable.
[GOAL] Create a 24-second product film showing a craftsperson assembling a portable desk lamp. [STAGE 1 — PARTS] Initial state: the lamp components are arranged separately on a clean wooden table. Primary event: the craftsperson connects the base, arm, and light head in that order. End state: the assembled lamp stands upright in the center; no loose parts remain in the hands. [STAGE 2 — ADJUSTMENT] Continue with the same person, table layout, and lamp structure. Primary event: the arm extends and the light head rotates toward an open sketchbook. End state: the beam lands on the center of the page and both hands leave the lamp. [STAGE 3 — HERO SHOT] The room becomes slightly darker while the lamp remains on. The camera makes a restrained push-in. Final state: the lamp is centered, fully assembled, and unchanged in shape. Maintain one lamp, stable materials, consistent hand count, continuous screen direction, and quiet workshop ambience throughout.
Use timestamps only when timing is part of the brief
Time range
Allocate a section of the runtime: “0–5s: establish the empty room.”
Exact moment
Reserve one beat: “At 8 seconds, the doors open.”
Relative timing
Link events: “Two seconds after the alarm, the lights turn red.”
Keep time ranges consecutive and non-overlapping. Treat them as pacing budgets, not frame-accurate edit points. Too little content leaves the model free to improvise; too much content encourages rushed cuts or skipped actions. For ordinary narrative work, stages should remain the default.
Separate generation, editing, and extension logic
A new video prompt describes what should be created. An editing prompt must protect a source video. An extension prompt must first connect to the source boundary. Treating these as different jobs makes the instruction much easier to follow.
For video editing: master, target, scope, preserve
Declare one source clip as the editing master. It controls the characters, scene, action, composition, camera, timing, and audio unless you explicitly change one of them. Then identify the target, narrow the time or region, and list the elements that must survive the edit.
[EDIT GOAL] Edit @Video 1. Between 4 and 7 seconds, change only the cool blue wall light to a warm amber light. [SOURCE MASTER] @Video 1 is the sole editing master for the character, room layout, action, framing, camera movement, event order, dialogue, and ambience. [EDIT SCOPE] Modify only the wall light and the surfaces it naturally illuminates. Let the skin and clothing respond subtly to the warmer light. [PRESERVE] Keep the character's identity, expression, position, motion, clothing, room structure, camera path, speech, and timing unchanged.
For extension: describe the boundary before the new event
A forward extension should begin from the source video’s final frame. A backward extension should finish at its first frame. Before adding a new action, describe the boundary state: subject pose, prop positions, scene geometry, camera direction, lighting, motion direction, and any continuous sound.
[SOURCE] @Video 1 is the video to extend forward. [BOUNDARY] Begin directly from the final frame of @Video 1. Preserve the subject's pose and direction, the bicycle's position, the street layout, the camera height, the afternoon light, and the current forward motion. [NEW ACTION] The cyclist exits the narrow street into an open plaza, slows beside a fountain, and looks toward the clock tower. The camera continues the existing tracking movement before easing into a wide reveal. [CONTINUITY] Keep the same cyclist, clothing, bicycle structure, travel direction, weather, and ambient city sound. Do not duplicate the cyclist or bicycle.
In input-led modes, aspect ratio or duration may be inherited from the source material. Use the generation interface for available settings and keep the prompt focused on content and continuity.
Advanced control without overloading the prompt
Keyframes, storyboards, blockouts, transition clips, and performance notes solve different problems. State which control each input owns instead of asking every asset to define the whole result.
First and last frames
Identify the opening and ending images separately, keep their aspect ratios aligned, and describe one continuous event that travels between them. Additional images should supplement identity, props, or materials without replacing either anchor composition.
Multiple keyframes and storyboards
Declare the image order first, then explain the state represented by each frame. Separate keyframe images are usually clearer than a single crowded grid. A storyboard controls sequence and key states; it is not a promise to reproduce every panel pixel for pixel.
Coarse and fine blockouts
Use a coarse blockout for motion paths, blocking, entrances, camera movement, cuts, lighting changes, and audio rhythm. Use a clean fine blockout when structure is already solved and the job is to re-render materials, colors, characters, scene detail, or style.
Image-to-video assembly
When turning a group of images into one finished video, specify each image's role, the intended order, how much motion to add, editing rhythm, visual packaging, and sound. “Make these images into a video” leaves too many production decisions undefined.
Seamless transitions
Define the outgoing clip, incoming clip, trigger action, camera direction, visual transformation, arrival composition, and audio handoff. Occlusion, object morphs, focus changes, and continuous pushes work best when corresponding shapes and motion directions are clear.
Direct performance with visible behavior
Words such as “tense,” “warm,” or “relieved” describe a mood, but they do not fully stage an actor’s performance. Add two to four observable cues: gaze, brow tension, mouth movement, breathing, shoulders, hands, or speech delivery. Use event-led stages only when the emotion changes more than once.
The performance moves from guarded concentration to quiet relief. When the mechanic hears the repaired engine start, the hands stop above the open hood and the eyes shift toward the dashboard. After the engine settles into a steady idle, the shoulders loosen, the breath releases slowly, and a small restrained smile appears. Keep the acting natural and understated. Preserve the character's position, clothing, and hand placement.
Seedance 2.5 prompt checklist
Before generating, read the prompt once as a production handoff. If a collaborator could not tell what changes, what stays fixed, or which reference controls which detail, the model will face the same ambiguity.
- The main subject and event are stated in the opening lines.
- Every reference has one named role and useful exclusions.
- Different views are identified as the same person or product.
- Long scenes are divided into stages with visible end states.
- Time ranges are consecutive and contain a realistic amount of action.
- Camera direction names the subject being followed or revealed.
- Dialogue includes its language, speaker, and delivery when needed.
- Edits define one master video, a narrow scope, and preserved content.
- Extensions describe the boundary frame before introducing new events.
- Constraints protect specific risks instead of becoming a generic negative list.
Seedance 2.5 prompt FAQ
Quick answers to the questions that usually appear between a first draft and a controllable generation.
What is the best structure for a Seedance 2.5 prompt?+
Start with the subject and the main action. Add the setting, visual treatment, camera direction, and sound only when they matter. For a longer scene, replace one dense paragraph with consecutive stages and a visible end state for each stage.
Should every reference be described in the prompt?+
Yes. Give each reference one clear job, such as identity, product structure, motion, camera movement, ambience, or music. Also state what should not be inherited when a reference contains an unwanted person, setting, or composition.
Do Seedance 2.5 prompts need timestamps?+
Not always. Stages are easier to maintain for most narratives. Use time ranges when pacing must be allocated, an exact timestamp for one critical beat, or relative timing when one action must happen after another.
How should I prompt a Seedance 2.5 video edit?+
Declare the source video as the editing master, identify the precise object, region, time range, or audio layer to change, then list what must remain untouched. Narrow edit scopes usually produce more predictable results.
What is the difference between a coarse and fine blockout?+
A coarse blockout mainly communicates motion, blocking, camera paths, and timing. A fine blockout already carries reliable structure, so the prompt can focus on re-rendering materials, color, characters, scene detail, and style.