Seedance·par ByteDanceReference to video
Seedance 2.5 Reference-to-Video
Condition on up to 50 reference assets — 30 images, 10 videos and 10 audio tracks — for up to 30 seconds of video. Reference them in the prompt with @Image1, @Video1, @Audio1 notation. Built for character consistency and music- or dialogue-driven scenes.
bytedance/seedance-2.5/reference-to-videoParamètres
| Paramètre | Type | Par défaut | Plage ou options |
|---|---|---|---|
duration durationVideo duration in seconds (4-30). | select | 5 | 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 |
resolution resolutionNative render resolution. 480p costs 0.465x and 1080p 1.968x of 720p. 480p and 720p are delivered as H.264; 1080p is delivered as H.265/HEVC. | select | 720p | 480p, 720p, 1080p |
output format output_formatOutput container. | select | mp4 | mp4, mov |
generate audio generate_audioGenerate a synchronized audio track. Does not affect price. | select | true | true, false |
ratio ratioAspect ratio. 'adaptive' lets the model choose. 21:9 costs ~0.5% more. | select | adaptive | 16:9, 4:3, 1:1, 3:4, 9:16, 21:9, adaptive |
return last frame return_last_frameAlso return the final frame as a separate still image. | select | false | true, false |
watermark watermarkAdd a watermark. | select | false | true, false |
Entrées
Fichiers que ce modèle accepte en plus du prompt.
| Entrée | Accepte | Fichiers max. | |
|---|---|---|---|
Reference Images reference_images | image | 30 | Facultatif |
Reference Videos reference_videos | video | 10 | Facultatif |
Reference Audio reference_audios | audio | 10 | Facultatif |
Guide de prompt pour Seedance
TARGET MODEL: Seedance 2.5 (ByteDance — native joint audio-video). The prompt is a four-part STRUCTURE written with a director's mind, not a one-liner. This is NOT Seedance 2.0: several of that model's laws are inverted here.
Règles à respecter
- STRUCTURE, in this order: (1) asset-reference line — number every supplied image/video/audio in upload order and state each one's job; (2) one-line summary — subject + place + event + genre/style + any special camera move; (3) beat body — split by integer-second timestamp OR "Shot N", each beat carrying content, camera, action, dialogue and SFX; (4) closing block — the global carry-through (camera position, environment, sound, atmosphere). Asset mapping first, global look and sound last. Single-beat fallback: subject + action detail + scene + light and color + camera + style + constraints.
- TIMESTAMPS WORK HERE and are integer-second. Three sanctioned forms: interval ("0-3 seconds… 3-7 seconds…", continuous, no gaps), point-in-time ("at the 5-second mark"), relative ("after 3 seconds"). NEVER convert a user's timecodes into "Shot N" ordering — that is the previous generation's rule and it is dead here. Never use a timestamp for frequency ("shake 3x per second"). Sub-second is not addressable, and a range is a time budget, not a frame-accurate edit point — never promise an exact cut.
- Audio is co-generated and on by default, so write it as CONTENT in the typed channels, with the character widths exactly as shown: music in fullwidth (…), sound effects in ASCII <…>, dialogue in ASCII {…}, on-screen subtitles in lenticular 【…】. Ambient beds ride in plain prose. Silence means OMITTING the () channel — never a promise or a toggle.
- Dialogue formula: language + regional variety or accent + delivery style + speaker + {line}. Any language other than Chinese or English MUST be tagged (says in Japanese {こんにちは}), and English must be tagged "in English" or the line tends to come back spoken in Chinese. One language per line, proper nouns excepted; keep lines short. If speech must land on the first frame, write "already speaking as the shot opens — first word on the very first frame, no silent lead-in", or the line is clipped at the duration mark.
- ONE camera move per shot. Verbatim: "Do not require push, pull, pan, and move at the same time, as this will increase image instability." Stack moves ACROSS beats, never inside one, and keep the camera in its own clause so its motion is never read as the subject's.
- Use the documented camera vocabulary, not synonyms. Shot size: extreme wide shot · wide shot · medium shot · medium close-up · close-up. Movement: push in · pull out · pan · track · follow · orbit · dive · pull back · tilt up · handheld shake. Angle: low angle · overhead shot · first-person perspective. Technique: one-shot/long take · Hitchcock (dolly) zoom · aerial perspective · FPV · bullet time · handheld shot · speed ramp. Anything outside that list must be written as [term + a plain description of what it does] — ByteDance's own example: "rack focus: the focus shifts smoothly; the trees that were clear in the foreground become blurred, while the character in the background gradually becomes clear." Drone becomes aerial perspective, FPV or dive; POV becomes first-person perspective.
- Spell framing as intent, never as shorthand: wide / normal / long, shallow / deep focus. No focal lengths, no f-numbers, no MCU/ELS/XCU/WS — the model buckets, and its own documentation calls 45mm a "wide-angle lens". Write "slight Dutch angle", never "35 degree Dutch angle". "35mm film look" and "film grain" are permitted as texture words only.
- Negation is permitted ONLY for subtitles and audio ("no subtitles", "no bgm"). Every VISUAL exclusion must be rewritten as a positive end-state: not "no cape" but "the shoulder line is clean and unbroken".
- Idiom is inert; physical description is not. Replace abstract emotion with externalized body detail — "eating with relish" becomes "a satisfied smile on his face, eating in big mouthfuls".
- Never author on-screen text and suppress text in the same shot — suppression is probabilistic, and a blanket "no text" tail kills the signage the user asked for.
- Resolution, aspect ratio, duration and the audio toggle are parameters — never write them into the prose. This model renders 480p and 720p only, so never claim 1080p or 4K in words.
- Length: about 60-100 words for a single beat, 150-250 for a timed multi-shot. Hard ceiling 500 Chinese characters / 1000 English words.
Ce qui fonctionne le mieux
- Register: "a good prompt is not descriptive copy, it is an engineering instruction." Mood words alone give the model nothing — "cinematic" is a documented failure when it stands in for blocking the shot, and valid only as a texture qualifier ("cinematic texture", "cinematic color grading").
- Lighting is the strongest axis: name a motivated source and its direction — "sunlight filters through the trees from the upper left". Documented terms: golden hour · blue hour · soft lighting · natural lighting and shadows · high contrast. Lighting is also referenceable ("refer to Image 1 for lighting and filters").
- Bind every visible subject to an asset and say WHAT to take from it, with an exclusion guard: "@Image 1 defines the artist's facial features, hairstyle and dark green apron — only her face, hair and clothing, never its background, never its vehicle, never its lighting." Unbound subjects duplicate, slot one leaks its whole frame, and the model reads sex from the reference face rather than from pronouns. Cap head-counts in the closing block.
- A transition needs both its trigger point and its method: "at the 5-second mark, the camera transitions leftward using a left wipe combined with a natural dissolve".
- Prioritise slow, coherent motion and always qualify pace — magnitude is unstable unqualified, and fast action morphs. Segment a long action into beats instead of writing one continuous take.
- Style by name lands for animation houses and film titles (Hayao Miyazaki, Makoto Shinkai, Disney 2D). Pair a live-action director's name with the concrete grammar it stands for; the name alone is unverified.
- Multi-shot is the hybrid header ByteDance's own examples use: "Shot 1 (0-3s): [extreme wide shot, ultra-low camera position…]". Give each range enough plot to fill it and no more — too little and the model improvises, too much and beats get dropped.
Issu du manuel LUVI pour cette famille : les mêmes règles qu'appliquent le réécriveur de l'espace de travail et le moteur de prompt MCP.