FLUX·by Black Forest LabsVideo extend

FLUX 3 Video (Extend)

FLUX 3 Video: continues an existing clip from its final frames, without a cut. The input video must be under 15 seconds and 50 MB; billing is on the GENERATED duration.

Open in workspaceblack-forest-labs/flux-3/extend-video

Parameters

ParameterTypeDefaultRange or options
Duration (s)
duration
select55, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20
Resolution
resolution
select720p720p, 1080p
Generate Audio
generate_audio
selecttruetrue, false
Safety Tolerance
safety_tolerance
slider20 – 4

Inputs

Files this model takes in addition to the prompt.

InputAcceptsMax files
Source Video
video_url
video1Required

Prompting guide for FLUX

TARGET MODEL: FLUX 3 Video (Black Forest Labs — native joint audio-video, preview model). One pass renders the picture AND its sound together: dialogue with lip-sync, ambience, effects and music, across 13+ languages including Turkish. The control surface is thin — no negative_prompt, no cfg/guidance, no weights, no seed, and nothing rewrites your prompt. What you write is what is read.

Rules that must hold

  • ONE framing term + ONE movement term + ONE subject action per sentence. Verbatim from the vendor: "low tracking shot" is clear, "low aerial handheld orbit push-in" usually is not. The discipline is on CAMERA vocabulary, not on subject vocabulary — camera and subject may both move.
  • Order: subject and action FIRST, then framing, movement and style around them. The camera phrase is a short prefix, never a paragraph. Do not front-load the way the FLUX image guides do — that rule is not restated for FLUX 3.
  • STATE THE MOTION MAGNITUDE. It is a documented slot, not an inference: slow, abrupt, weightless, chaotic, precise, cinematic, documentary. The vendor's own examples always prepend it.
  • AUDIO IS WRITTEN IN, in four layers, and you do not need all four — one or two is usually right, because a busy mix is a documented failure. Speech (who, the exact words, how) · Ambience (the sound of the place) · Effects (tied to a VISIBLE action: "the mug clicks against the saucer when he sets it down") · Music (style plus where it sits in the mix).
  • DIALOGUE IS QUOTED AND GIVEN A VISIBLE SPEAKER. A quoted line with no visible speaker and no explicit voiceover/narration cue MAY RENDER AS ON-SCREEN TEXT instead of speech — this is the single best-documented failure of the model. A visible speaker gives it a face to lip-sync; an off-screen line needs the word "voiceover" or "narration".
  • NEVER ASK FOR SILENCE IN PROSE. "quiet room tone" collapses into static or dead air. Name the sound you want instead. Full silence is the generate_audio parameter, not a sentence. Partial silence is fine as content: "for the final two seconds, only rain against the window".
  • A MIX NEEDS A FOREGROUND LAYER OR IT DRIFTS TOWARD INAUDIBLE. Measured across ten clips: every soundscape built only from ambience and effects landed between -36 and -61 LUFS — the bottom of that range is effectively silence — while every clip carrying dialogue or a music bed sat between -18 and -32. Naming the sounds is necessary but not sufficient. When the audio matters and the scene has no speech, give it a music cue and say where it sits in the mix.
  • Label each spoken line's language, and keep lines SHORT — a short line in a longer clip is safer than copy written to fill every second. Multi-speaker: attribute by visible ROLE, never by name, with short separate turns and no overlap.
  • Lens and framing are CATEGORY WORDS, not numbers. No f-stops, no film stocks, no camera-body names, no hex colours — all of that is image-guide vocabulary with zero documented effect here. "Wide angle (24mm)" is the ONLY focal length in the entire video documentation, and the number is redundant with its bucket. Depth of field goes by named technique: "shallow (sharp on subject, blurred background)". Colour goes in prose: "warm backlight with soft rim — amber, cream, walnut".
  • Use real grip terms, and the right ones: the vendor honours pedestal, crane/boom, trucking, dolly in, arc shot, camera roll, locked-on and steadicam follow as DISTINCT moves. "Push-in" and "pull-back" are not enum terms — write "dolly in" or "push through". Bare "zoom" is not documented — only "dolly zoom".
  • Concrete nouns and verbs the camera can actually see. Vague adjectives leave the result to chance; you are directing a scene, not describing a collection of objects.
  • Length is not the goal. Short prompts leave details to chance; over-stuffing makes MOTION LESS COHERENT. About 330 words is the vendor's own anchor for a 10-second, 4-shot sequence.
  • Negation is first-party practice here despite there being no negative_prompt field — but ONLY for audio layers and screen furniture: "No on-screen text, no subtitles." · "no announcer delivery" · "no music, no second voice". Never negate into silence.
  • Kill ad-copy voice. If a read sounds like an advert, the fix is a person plus a recording setup plus a social situation, not praise words: "close and dry, warm mid register, lightly amused".

What works best

  • MULTI-SHOT INSIDE ONE GENERATION is this model's headline capability, and the only place continuity is model-enforced rather than prompt-hoped — there is no seed and no style reference, so prefer one multi-shot generation over stitched clips. HARD CUT is the documented cut token: "SHOT ONE: ... HARD CUT. SHOT TWO: ...". Character, look and continuity hold across the cuts with a single audio bed carrying through.
  • Continuity is carried by a VERBATIM-IDENTICAL SUBJECT PARAGRAPH repeated in every shot, plus one global style-and-colour block.
  • Beat density: two or three beats per 5 seconds. The vendor's 10-second example is 4 shots of 2.5s.
  • TIMECODES ARE ORDERING, NOT MARKS. Measured on this surface: a single HARD CUT written at 5.0s of a 10s clip rendered as cuts at 1.46s and 7.71s, while shot order, identity and dialogue placement all held. Write beats as a timeline when the action has to land in sequence — never promise a frame-accurate in-point.
  • Three registers, pick by need: the one-liner default "[camera] shot of [subject] [action] in [environment]. [supporting visual and motion details]"; labelled fields when tuning levers (Camera shot / Subject + action / Depth of field / Lighting + palette / Motion / Style); the full schema for multi-shot (Core summary · Scene · Subject description · Dynamic narrative · Audio · Style & colour).
  • Style is prose only. Vetted terms from the vendor's own examples: "cinematic naturalism", "gritty documentary realism", "epic western realism", "fine grain adds gritty, tactile realism". There is no raw mode — ask for realism AFFIRMATIVELY (grain, pores, unretouched texture); you cannot ask for plastic sheen to be removed.
  • A director's name only when paired with the mechanical description that does the work ("staged like a Wes Anderson set — uniformed staff in neat formation, deadpan symmetry"). The name alone underperforms.
  • NO TAKE IS REPEATABLE on this surface: there is no seed and no draft loop, and two identical submits differ completely. Never write a rewrite that implies a variation of a previous render.

From LUVI's own manual for this family, the same rules the workspace rewriter and the MCP prompt engine apply.

Other models in this family

FLUX 3 Video (Extend): price, parameters and prompting guide · LUVI Creator