Veo·by GoogleImage to video

Veo 3.1 Lite Image-to-Video

Animate an image into video economically with Veo 3.1 Lite. The cost-effective option for scalable workflows.

Open in workspacegoogle/veo3.1-lite/image-to-video

Parameters

ParameterTypeDefaultRange or options
duration
duration

Length of the generated video in seconds. Must be 8 when using 1080p resolution.

select84, 6, 8
aspect ratio
aspect_ratio

The aspect ratio of the generated video.

select16:916:9, 9:16
resolution
resolution

Output resolution. 1080p only supports 8s duration and costs more per second. Veo 3.1 Lite does not support 4K.

select720p720p, 1080p
seed
seed

Random seed. Does not guarantee determinism but may improve repeatability.

number

Inputs

Files this model takes in addition to the prompt.

InputAcceptsMax files
image
image
image1Required

Prompting guide for Veo

TARGET MODEL: Veo (Google DeepMind video — the shot arrives already scored: picture AND audio generate in the same pass, lip-synced dialogue, SFX, ambience, music). Write a described shot, not tags — a camera, a subject doing something, a world, a look.

Rules that must hold

  • Slot order (Google's published, rewarded order): Cinematography → Subject → Action → Context → Style & Ambiance. Lead with the camera, name the subject, give it a verb, place it, then set the stock and the light.
  • Front-load the subject and its action — the first clause sets adherence.
  • ONE clean beat per clip. A beat is a single action that completes inside the clip. For more than one move, timestamp the beats — "[00:00-00:02] …", "[00:02-00:04] …" — one shot + one sound cue per bracket. Never cram a scene into one unbroken description.
  • Camera moves by their PROPER classical names — "dolly shot", "tracking shot", "crane shot", "aerial view", "slow pan", "tilt", "arc", "POV shot"; shot size and angle by name ("wide shot", "close-up", "extreme close-up", "low angle", "two-shot"); optics by name ("shallow depth of field", "wide-angle lens", "macro lens", "deep focus"). One move per beat — stacked moves muddy all of them.
  • In-prose exclusions are DESCRIBED ABSENCES, never prohibitions: "a desolate landscape with no buildings or roads" describes emptiness; never write bare "no X" / "don't" commands in the shot.
  • Audio is written INTO the shot: the exact spoken line in quotes (it will be lip-synced), attributed to a distinct speaker — 'A woman says, "We have to leave now."'; sound effects cued by name — "SFX: thunder cracks in the distance"; room tone set — "Ambient noise: the quiet hum of a starship bridge"; music named — "soft piano score builds under the scene"; silence is directable — state it and specify ambience only ("no music, only the wind").

What works best

  • Use DeepMind's seven nameable elements as a checklist so the model isn't guessing defaults: shot & motion, style, lighting, character, location, action, dialogue.
  • Resolve the subject to matter (age, wardrobe, material, wear), then give it one clear action that UNFOLDS over the clip — describe how the movement progresses, not a frozen pose.
  • Physics is a prompt: name weight, contact, momentum. Vague verbs drift.
  • Lighting = source + direction + quality + temperature, plus time-of-day/weather ("golden hour", "overcast", "neon night"). Named setups land as intent.
  • One dominant look — genre, era, film-stock reference, animation style — as a clause, never a stack. Ambiance words ("tense", "serene", "dreamlike") ride the tail; they set the grade, not the subject.
  • Keep spoken lines short enough for the clip — roughly a sentence or two per 8 seconds. Bind each line to the action that motivates it: describe the beat, then give the line.
  • Continuity across shots: repeat the character/wardrobe/light description and keep camera language consistent.

From LUVI's own manual for this family, the same rules the workspace rewriter and the MCP prompt engine apply.

Other models in this family