Seedream·by ByteDanceImage edit
Seedream v5.0 Lite Edit Sequential
ByteDance next-generation image editing model with batch generation support. Edit multiple images while preserving facial features and details.
bytedance/seedream-v5.0-lite/edit-sequentialParameters
| Parameter | Type | Default | Range or options |
|---|---|---|---|
output format output_formatOutput image file format. Seedream v5.0 Lite supports both JPEG and PNG output. | select | jpeg | jpeg, png |
max images max_imagesThe maximum number of images to generate (1-15). The total of input reference images plus generated images must not exceed 15. | slider | 1 | 1 – 15 |
size sizeOutput image size in 'WIDTH*HEIGHT' pixels. 16 preset resolutions at 2K and 3K tiers. Total pixel range: [3,686,400, 10,404,496]. Aspect ratio range: [1/16, 16]. | select | 2048*2048 | 2048*2048, 2304*1728, 1728*2304, 2848*1600, 1600*2848, 2496*1664, 1664*2496, 3136*1344, 3072*3072, 3456*2592, 2592*3456, 4096*2304, 2304*4096, 2496*3744, 3744*2496, 4704*2016 |
Inputs
Files this model takes in addition to the prompt.
| Input | Accepts | Max files | |
|---|---|---|---|
images images | image | 14 | Required |
Prompting guide for Seedream
TARGET MODEL: Seedream (ByteDance image — generation and editing are one unified model). The prompt is the whole instrument: write a finished CAPTION in prose, not tags.
Rules that must hold
- Lead-weighting is the exploitable behavior: concepts mentioned EARLIER weigh more. The first clause is the hero — put what must survive first, negotiable detail last.
- The formula, left-to-right as priority order: Subject → Scene/Setting → Composition & Framing → Style → Lighting & Atmosphere → Camera/Lens.
- Length: 30-100 words is the sweet spot (hard ceiling ~600); more words only when they add a decision — padding dilutes the lead.
- Kill noise tokens: "masterpiece, best quality, 8K, ultra-detailed, trending" — delete them.
- Aspect ratio is a parameter, never prose — do not write "16:9" or frame ratios into the caption.
- On-image copy goes in "double quotes", in the TARGET SCRIPT (actual glyphs — "静かな庭", never romanized). Unquoted text renders as the idea of a sign, not the words.
- Editing REGENERATES the entire frame — "keep X unchanged" is guidance to imitate, never a pixel lock. Never promise preservation of a real capture.
What works best
- Every noun carries a modifier — resolve the subject to something a camera could find: age, hair, expression, wardrobe, material. "A woman" is a stock face.
- State shot scale ("medium shot", "close-up", "overhead perspective"); stronger still, name the focal point inside the scene ("his hands are the focal point") and put the camera in space ("taken from 3000 feet").
- Name geometry ("symmetrical composition", "rule of thirds", "foreground detail with blurred background") and anchor every element with prepositions ("on the left side", "between them", "arranged left to right") — ambiguity is the only thing that drifts.
- Photoreal fork: "photorealistic rendering / ultra-detailed / commercial quality" produces a clean CGI look — right for products, wrong for faces. For a photograph, describe a CAPTURE plus a genre ("editorial photography", "documentary", "street photography").
- Camera report as observation: focal length as grammar (24-35mm environmental, 50-85mm portrait, 100-200mm compression); aperture by its effect ("the bokeh turns harbor lights into perfect circles"); shutter as motion ("1/8000 freezes", "30-second exposure trails the lights").
- Pair a film stock with a color line: "Kodak Portra 400, warm color cast, visible film grain" beats the stock name alone; push language works ("expired Portra 800, pushed two stops, halation, lifted blacks").
- Motivated light: direction + quality + color, optionally a named recipe ("warm tungsten key from frame-left, cool cyan neon rim"; "chiaroscuro", "Rembrandt", "Vermeer lighting").
- Fight plastic skin texture-forward and positive: "weathered face", "sun-kissed skin", "visible pores and fine lines", "matte skin", "no beauty retouching", "visible film grain" — never the phrase "realistic skin".
- Style: ONE named anchor, set twice as bookends — a register word opens the caption, a finishing tag closes it. Director and stock references are native ("in the manner of Deakins", "Wes Anderson palette", "film noir"). Blend at most one movement + one technique; reach for a style's keyword cluster ("soft edges, translucent layers, paper texture") when the bare name renders timid.
- Typography is the signature strength: give quoted copy a surface and a treatment ("neon sign reading …", "engraved brass plate", "handwritten chalk"); spec type like a designer — character, weight, case, color, and a grid position ("top-left", "along the bottom edge"). Keep each line ~3-5 words; numerals, currency and dates are copy too. Dense body copy or fine print (when it is NOT user-quoted verbatim copy): cut the word count and describe it by role rather than transcribing it. Posters: name the medium, then stack each line top-to-bottom with copy + size word + anchor.
- Sequential sets: trigger with "series", "set", "sequence", a numeric count, or an explicit grid ("2x2 storyboard grid", "three-view drawing"); copy the character description VERBATIM between panels; declare each frame's copy, shot scale and vantage so the set reads as one designed system.
From LUVI's own manual for this family, the same rules the workspace rewriter and the MCP prompt engine apply.