How to prompt FLUX 3 Video—and why your dialogue came back as on-screen text
FLUX 3 Video renders picture and sound in one pass, so the audio is written into the prompt rather than switched on. Quote every line and give it a visible speaker, or it can appear as text burned into the frame. On LUVI, clips start at 1,870 credits for five seconds at 720p.

FLUX 3 Video is Black Forest Labs' joint audio-video model: one pass renders the picture and its sound together, including lip-synced dialogue. On LUVI it runs as text-to-video, image-to-video, keyframes-to-video and extend-video, at 720p or 1080p, from five to twenty seconds, starting at 1,870 credits for a five-second clip at 720p. The vendor's page tells you what the model can do. This post is about how it reads a prompt—starting with the mistake that costs people the most renders.
Why did my dialogue come back as on-screen text?
A quoted line with no visible speaker can render as text burned into the frame instead of speech. This is the best-documented failure of the model, and it follows from a capability Black Forest Labs advertises: FLUX 3 can "render typography as a natural part of the scene." The model is genuinely good at putting words on screen, so an unattributed quotation is ambiguous—it can read as a line to be spoken or a line to be shown.
Our manual for the family states the rule plainly: dialogue is quoted and given a visible speaker. In practice that means three things.
- Put a face in the shot. A visible speaker gives the model a mouth to lip-sync. "She says, ..." beats a floating quotation every time.
- Say the word "voiceover" or "narration" when the line is deliberately off-screen. Without it, an off-screen quotation is the exact ambiguous case.
- Label the language and keep the line short. Black Forest Labs lists lip-sync support for "English (various dialects), Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi, Punjabi and more." A short line in a longer clip is safer than copy written to fill every second.
If you want to be certain nothing appears on screen, negation works here for screen furniture: "No on-screen text, no subtitles."
How do you write the sound?
Sound on FLUX 3 Video is content you write, not a setting you turn up, and it comes in four layers: speech, ambience, effects tied to a visible action, and music with a note about where it sits in the mix. You rarely want all four—a busy mix is a documented failure mode. One or two layers is usually right.
There is a measured catch. Across ten clips we generated on this family, every soundscape built only from ambience and effects landed between -36 and -61 LUFS; the bottom of that range is effectively silence. Every clip carrying dialogue or a music bed sat between -18 and -32 LUFS. Naming your sounds is necessary but not sufficient: a mix with no foreground layer drifts toward inaudible. When the audio matters and nobody speaks, give the clip a music cue and say where it sits.
Medium shot of a barista in a green apron sliding a paper cup across the counter toward the customer in front of her, slow dolly in, warm morning light through the window. She says in English, "Careful, it's very hot." Ambience: low cafe chatter and the hiss of a steam wand. The cup clicks against the saucer when she sets it down. No on-screen text, no subtitles.
Can you ask for a silent clip?
Not in prose. Asking for "quiet room tone" or "complete silence" can collapse into static or dead air, because you are still asking the audio model to generate something. Full silence is the generate_audio parameter, which you switch off in the Workspace. Partial silence is fine as content, because it is a description of a mix rather than an absence of one: "for the final two seconds, only rain against the window."
On LUVI, switching the audio off does not change what you pay. A five-second clip at 720p estimates at 1,870 credits with generate_audio on and 1,870 credits with it off (checked September 22, 2026). That is worth knowing because it is not the rule everywhere—on Veo 3.1 the same parameter is off by default and the model's own description says it "increases the per-second cost."
How do you get several shots out of one clip?
Multi-shot inside a single generation is this model's headline capability, and the cut token is HARD CUT. Write "SHOT ONE: ... HARD CUT. SHOT TWO: ...", repeat the subject description verbatim in every shot, and name one audio bed once for the whole run.
This matters more than it sounds, because FLUX 3 Video has no seed and no draft loop: two identical submissions come back completely different, and there is no style reference to carry a character across separate renders. One multi-shot generation is the only place continuity is enforced by the model rather than hoped for by the prompt.
Timecodes are ordering, not marks. In our test, a single HARD CUT written at 5.0 s of a ten-second clip rendered as cuts at 1.46 s and 7.71 s—while shot order, identity and dialogue placement all held. Write beats as a timeline when the sequence matters; never promise yourself a frame-accurate cut point.
SHOT ONE: A fisherman in a yellow oilskin coat hauls a dripping net over the rail of a small wooden boat, slow tracking shot, gray dawn light. HARD CUT. SHOT TWO: A fisherman in a yellow oilskin coat drinks from a steel flask in the wheelhouse, locked-on medium shot, gray dawn light. Music: a slow low cello bed under the whole clip. Ambience: gulls, wind, and water knocking against the hull.
Which habits from image models don't transfer?
The photography vocabulary that works on FLUX image models has no documented effect on FLUX 3 Video, and the parameters are not words.
- No f-stops, film stocks, camera bodies or hex colors. Write the technique instead: "shallow (sharp on subject, blurred background)", "fine grain adds gritty, tactile realism", "amber, cream, walnut".
- No weights and no cfg.
(melancholy:1.4)is not read; the parentheses are noise around a word the model would have taken anyway. - No Negative Prompt field. Exclusions are allowed, but only for audio layers and screen furniture: "no music, no second voice", "No on-screen text, no subtitles." Never negate your way into silence.
- Aspect ratio, resolution and duration are parameters. Writing "16:9", "1080p" or "10 seconds" into the prompt does nothing.
- Use the grip terms the model knows. "Dolly in", "push through", "dolly zoom", "trucking", "arc shot", "pedestal", "locked-on". Bare "zoom" and "push-in" are not documented terms. One framing term, one movement term and one subject action per sentence: "low tracking shot" is clear, "low aerial handheld orbit push-in" usually is not.
How we checked
- Models:
black-forest-labs/flux-3/text-to-videoandblack-forest-labs/flux-3/image-to-video, plus the extend and keyframes modes for prices. - Settings: five to twenty seconds, 720p and 1080p,
generate_audioon and off. - Date and what was read: September 22, 2026. Model schemas and credit estimates from LUVI, the prompting guide LUVI runs for this family, and the Black Forest Labs pages linked below.
- Results: 1,870 credits for five seconds at 720p with audio either on or off; 6,380 credits for ten seconds at 1080p; 7,480 credits for twenty seconds at 720p. The LUFS range and the cut-placement figure come from an earlier set of ten clips generated on this family while the guide was being written.
- Drawbacks: No outputs were generated for this post, so it makes no claim about output quality. Prices move when a provider reprices a model, so re-check the estimate shown in the Workspace before a long clip.
In LUVI
Pick FLUX 3 Video in the Workspace, write the prompt, and press the feather button next to it before you generate. That is LuviTransLex, the free prompt check that applies this model's prompting guide and rewrites what it can.
It is worth running even on a prompt that looks fine. A line like A man sits alone at a kitchen table, 16:9, 1080p, shot on Kodak Portra at f/1.8, (melancholy:1.4), quiet room tone, #1a1a1a shadows, 10 seconds reads as careful prompting and comes back with one rewrite and eight warnings: the weight dropped and the word kept, the film stock and f-stop flagged as image vocabulary, the hex color flagged, the ratio, resolution and duration flagged as parameters, and the silence request flagged. Checks and rewrites cost nothing.
FLUX 3 Video has no Negative Prompt field, which is why the field is missing from its panel—when a model has one and when it doesn't is a property of the model, not a setting. You can also drive the whole thing from Claude or ChatGPT through the LUVI connector.
New here? Create a LUVI account and the sign-up credits are enough to try a short clip.
Sources
- FLUX 3 Video, Part 1: Generation — Black Forest Labs, product blog, August 4, 2026. Accessed September 22, 2026.
- FLUX 3 documentation — Black Forest Labs, API documentation. Accessed September 22, 2026.
- FLUX 3: One Multimodal Model — Black Forest Labs, model page. Accessed September 22, 2026.
More from Guides
- How to prompt Hailuo: fifteen camera commands, written in square bracketsHailuo takes its camera direction as a bracketed command from a closed list of fifteen, placed inline where the move happens—[Push in], [Tracking shot], [Static shot]. It is also picture only, so every word you spend on sound is wasted. Six seconds start at 380 credits.
- How to prompt Wan 3.0—name the camera in every shot, or it cuts for youWan 3.0 reads an additive formula—entity, scene, motion, then the look, then the sound—and it renders audio in the same pass. Leave the camera unnamed and it will cut inside a clip you wanted as one take. Five seconds runs from 400 credits at 480p.
- Hunyuan3D doesn't read your prompt: on image-to-3D, the picture is the whole briefOn Hunyuan3D's image-to-3D modes the text field is not read at all—the model's own contract says it accepts no prompt. Every bit of control you have is in how you prepare the input image. Meshes start at 600 credits on Rapid and 1,000 on Pro.