by GoogleVideo

Veo

7 models · from 800 credits

Models in the Veo family

ModelModePriceAspect ratios
Veo 3.1 Fast Image-to-VideoImage to videofrom 1,280 credits16:9 · 9:16Open
Veo 3.1 Fast Text-to-VideoText to videofrom 1,280 credits16:9 · 9:16Open
Veo 3.1 Image-to-VideoImage to videofrom 3,200 credits16:9 · 9:16Open
Veo 3.1 Lite Image-to-VideoImage to videofrom 800 credits16:9 · 9:16Open
Veo 3.1 Lite Start-End FrameStart and end framefrom 800 credits16:9 · 9:16Open
Veo 3.1 Lite Text-to-VideoText to videofrom 800 credits16:9 · 9:16Open
Veo 3.1 Reference-to-VideoReference to videofrom 3,200 creditsOpen

Prompting guide for Veo

TARGET MODEL: Veo (Google DeepMind video — the shot arrives already scored: picture AND audio generate in the same pass, lip-synced dialogue, SFX, ambience, music). Write a described shot, not tags — a camera, a subject doing something, a world, a look.

Rules that must hold

  • Slot order (Google's published, rewarded order): Cinematography → Subject → Action → Context → Style & Ambiance. Lead with the camera, name the subject, give it a verb, place it, then set the stock and the light.
  • Front-load the subject and its action — the first clause sets adherence.
  • ONE clean beat per clip. A beat is a single action that completes inside the clip. For more than one move, timestamp the beats — "[00:00-00:02] …", "[00:02-00:04] …" — one shot + one sound cue per bracket. Never cram a scene into one unbroken description.
  • Camera moves by their PROPER classical names — "dolly shot", "tracking shot", "crane shot", "aerial view", "slow pan", "tilt", "arc", "POV shot"; shot size and angle by name ("wide shot", "close-up", "extreme close-up", "low angle", "two-shot"); optics by name ("shallow depth of field", "wide-angle lens", "macro lens", "deep focus"). One move per beat — stacked moves muddy all of them.
  • In-prose exclusions are DESCRIBED ABSENCES, never prohibitions: "a desolate landscape with no buildings or roads" describes emptiness; never write bare "no X" / "don't" commands in the shot.
  • Audio is written INTO the shot: the exact spoken line in quotes (it will be lip-synced), attributed to a distinct speaker — 'A woman says, "We have to leave now."'; sound effects cued by name — "SFX: thunder cracks in the distance"; room tone set — "Ambient noise: the quiet hum of a starship bridge"; music named — "soft piano score builds under the scene"; silence is directable — state it and specify ambience only ("no music, only the wind").

What works best

  • Use DeepMind's seven nameable elements as a checklist so the model isn't guessing defaults: shot & motion, style, lighting, character, location, action, dialogue.
  • Resolve the subject to matter (age, wardrobe, material, wear), then give it one clear action that UNFOLDS over the clip — describe how the movement progresses, not a frozen pose.
  • Physics is a prompt: name weight, contact, momentum. Vague verbs drift.
  • Lighting = source + direction + quality + temperature, plus time-of-day/weather ("golden hour", "overcast", "neon night"). Named setups land as intent.
  • One dominant look — genre, era, film-stock reference, animation style — as a clause, never a stack. Ambiance words ("tense", "serene", "dreamlike") ride the tail; they set the grade, not the subject.
  • Keep spoken lines short enough for the clip — roughly a sentence or two per 8 seconds. Bind each line to the action that motivates it: describe the beat, then give the line.
  • Continuity across shots: repeat the character/wardrobe/light description and keep camera language consistent.

From LUVI's own manual for this family, the same rules the workspace rewriter and the MCP prompt engine apply.