How to prompt Veo 3.1 — and why your clip came back silent

Write Veo 3.1 prompts camera first: cinematography, subject, action, context, style. Put dialogue, effects and ambience into the shot, and switch on Audio Generation, which is off by default on LUVI. Veo 3.1 Lite has no audio. Clips start at 800 credits for eight seconds at 720p.

Published:

A cinema lens iris in the foreground frames a silhouette speaking, warm sound ripples spreading out, and an amber toggle switch glowing on.

Veo 3.1 is Google DeepMind's video model, and one prompt can give you picture and sound together: dialogue, effects, ambience and music. On LUVI the Veo family has seven models, from 800 credits for an eight-second 720p clip on Veo 3.1 Lite (prices as of 16 September 2026). Two things decide whether a Veo 3.1 prompt works: the order you write it in, and whether you asked for sound at all. This guide covers both, from the prompting guide we run inside LUVI and from Google's own documentation.

Why did my Veo 3.1 clip come back silent?

A Veo 3.1 clip on LUVI is silent unless you switch on Audio Generation before you generate. The switch is off by default on every Veo model that has one, because sound changes the price, and writing a line of dialogue into the prompt does not turn it on. On an eight-second 720p clip:

  • Veo 3.1 Fast (text-to-video and image-to-video) costs 1,280 credits silent and 1,536 credits with audio.
  • Veo 3.1 (image-to-video and reference-to-video) costs 3,200 credits silent and 6,400 credits with audio.
  • Veo 3.1 Lite has no audio switch at all. Every Lite clip is silent, whatever the prompt says.

A second rule catches people out. When you give Veo 3.1 or Veo 3.1 Fast an end frame as well as a start frame, LUVI turns audio off for that run, because these models can't generate sound when an end frame is supplied. If the soundtrack matters more than the exact final frame, leave the end frame out.

What order should a Veo 3.1 prompt follow?

A Veo 3.1 prompt should lead with the camera: cinematography, then subject, action, context, and style and ambiance. Google publishes this five-part formula in its ultimate prompting guide for Veo 3.1, where cinematography means "the camera work and shot composition". Our prompting guide enforces the same order and adds one rule: the subject and its action belong in the first clause, because that clause sets how closely Veo follows the rest.

  • Cinematography. Name the move and the framing with classical terms: dolly shot, tracking shot, crane shot, slow pan, close-up, low angle. Use one move per beat, because stacked moves blur each other.
  • Subject. Resolve it to material detail: age, clothing, wear.
  • Action. Give one action that unfolds over the clip, described as movement rather than a pose. Weight and contact help; vague verbs drift.
  • Context. Name the place and a light source with its direction and colour.
  • Style and ambiance. Pick one dominant look, such as a film stock or an era, and put mood words at the end.

Here is a single-beat prompt in that order, for Veo 3.1 Fast Text-to-Video with Audio Generation on:

Slow dolly shot, medium close-up: an elderly clockmaker in a worn leather apron lifts a brass gear toward the lamp and turns it between two fingers, squinting at its teeth. A cramped workshop at night, walls lined with ticking clocks, one warm desk lamp from the left and the rest in shadow. He says quietly, "There you are." SFX: a soft metallic click as the gear seats. Ambient noise: dozens of clocks ticking out of step. Shot on 35mm film, warm tungsten grade, calm and intimate.

How do I write dialogue and sound for Veo 3.1?

Veo 3.1 reads sound as part of the shot description, in three labelled kinds. Google's guide gives the pattern: the spoken line in quotation marks after a named speaker, sound effects after SFX:, and the background after Ambient noise:. Music is described in plain words, such as a soft piano score building under the scene.

  • Keep lines short. A sentence or two fits an eight-second clip. Describe the action first, then give the line it motivates.
  • Give each line one distinct speaker. Veo 3.1 lip-syncs the line to the character you attribute it to.
  • Quotation marks or a colon. Google's own documents differ here. The Cloud blog quotes the line, while the Vertex AI video prompt guide writes it after a colon: The seasoned detective says: Your story has holes. Our prompting guide uses quotation marks.
  • Write the prompt in English. The Veo 3.1 model page on Google Cloud lists English as the only prompt language, for Veo 3.1 and Veo 3.1 Fast alike.
  • Direct silence instead of forbidding sound. For a quiet scene, say what should be heard: "no music, only the wind".

All of this only plays if Audio Generation is on.

How do I keep something out of a Veo 3.1 shot?

Veo 3.1 on LUVI has no negative prompt field, so an exclusion goes into the prompt as a description of what is there. Google's example is "a desolate landscape with no buildings or roads" in place of "no man-made structures". In its section on negative prompts, the Vertex AI guide lists "no walls" and "don't show walls" as wording to avoid. Describe the empty beach, not the missing people.

How do I fit more than one action into eight seconds?

Split the clip into timestamped beats, with one shot and one sound cue per bracket. Veo 3.1 clips on LUVI run 4, 6 or 8 seconds (reference-to-video runs 8 only), and one unbroken description of a whole scene tends to blur. Google calls this timestamp prompting and writes the brackets as [00:00-00:02]:

[00:00-00:04] Wide shot, slow pan across an empty beach at dawn, smooth wet sand with no footprints, low pink light from the horizon. Ambient noise: small waves folding onto the shore. [00:04-00:08] Close-up as a dog's paw presses the first print into the sand. SFX: a soft wet thud, then distant gulls.

Which Veo 3.1 model should I pick?

The right Veo 3.1 model depends on what you start from and whether you need sound. All of them output 16:9 or 9:16, at 720p by default.

Google now points Gemini API developers to Gemini Omni Flash as the default video model, and keeps Veo 3.1 for "scene extension, last-frame control, or integration with legacy pipelines", according to its video generation overview. If you don't need those, compare Gemini Omni 1.1 Flash on LUVI and our H3, Omni and Wan 3.0 comparison.

How we checked

  • What we read: the parameters of all seven Veo models on LUVI and LUVI's Veo prompting guide, on 16 September 2026.
  • Prices: LUVI's own estimate for eight-second 720p clips, with and without Audio Generation, on 16 September 2026. Longer settings, 1080p and 4K cost more.
  • Example prompts: checked against the Veo prompting guide. We didn't generate clips for this guide, so it makes no claim about output quality.

In LUVI

In the Workspace, pick a Veo model, set Audio Generation before you write, and draft the prompt in the order above. The feather button next to the prompt rewrites a draft into the Veo dialect for free; what a prompting guide is explains where those rules come from. If you work from Claude or ChatGPT, the LUVI connector returns each Veo model's exact parameters, including the audio switch. To try it, create a LUVI account.

Sources

More from Guides