Veo·de GoogleReference to video

Veo 3.1 Reference-to-Video

Generate video guided by 1-3 reference images while preserving character and style consistency across scenes with Veo 3.1.

Abrir en el espacio de trabajogoogle/veo3.1/reference-to-video

Parámetros

ParámetroTipoPor defectoRango u opciones
duration
duration

The duration of the generated video in seconds (reference-to-video supports 8s only).

select88
resolution
resolution

Video resolution. Higher resolutions increase the per-second cost.

select720p720p, 1080p, 4k
seed
seed

Random seed. Does not guarantee determinism but may improve repeatability.

number
generate audio
generate_audio

Whether to generate synchronized audio (dialogue, SFX, ambience). Increases the per-second cost.

selectfalsetrue, false

Entradas

Archivos que este modelo recibe además del prompt.

EntradaAceptaMáx. de archivos
images
images
image3Obligatoria

Guía de prompts para Veo

TARGET MODEL: Veo (Google DeepMind video — the shot arrives already scored: picture AND audio generate in the same pass, lip-synced dialogue, SFX, ambience, music). Write a described shot, not tags — a camera, a subject doing something, a world, a look.

Reglas que deben cumplirse

  • Slot order (Google's published, rewarded order): Cinematography → Subject → Action → Context → Style & Ambiance. Lead with the camera, name the subject, give it a verb, place it, then set the stock and the light.
  • Front-load the subject and its action — the first clause sets adherence.
  • ONE clean beat per clip. A beat is a single action that completes inside the clip. For more than one move, timestamp the beats — "[00:00-00:02] …", "[00:02-00:04] …" — one shot + one sound cue per bracket. Never cram a scene into one unbroken description.
  • Camera moves by their PROPER classical names — "dolly shot", "tracking shot", "crane shot", "aerial view", "slow pan", "tilt", "arc", "POV shot"; shot size and angle by name ("wide shot", "close-up", "extreme close-up", "low angle", "two-shot"); optics by name ("shallow depth of field", "wide-angle lens", "macro lens", "deep focus"). One move per beat — stacked moves muddy all of them.
  • In-prose exclusions are DESCRIBED ABSENCES, never prohibitions: "a desolate landscape with no buildings or roads" describes emptiness; never write bare "no X" / "don't" commands in the shot.
  • Audio is written INTO the shot: the exact spoken line in quotes (it will be lip-synced), attributed to a distinct speaker — 'A woman says, "We have to leave now."'; sound effects cued by name — "SFX: thunder cracks in the distance"; room tone set — "Ambient noise: the quiet hum of a starship bridge"; music named — "soft piano score builds under the scene"; silence is directable — state it and specify ambience only ("no music, only the wind").

Lo que mejor funciona

  • Use DeepMind's seven nameable elements as a checklist so the model isn't guessing defaults: shot & motion, style, lighting, character, location, action, dialogue.
  • Resolve the subject to matter (age, wardrobe, material, wear), then give it one clear action that UNFOLDS over the clip — describe how the movement progresses, not a frozen pose.
  • Physics is a prompt: name weight, contact, momentum. Vague verbs drift.
  • Lighting = source + direction + quality + temperature, plus time-of-day/weather ("golden hour", "overcast", "neon night"). Named setups land as intent.
  • One dominant look — genre, era, film-stock reference, animation style — as a clause, never a stack. Ambiance words ("tense", "serene", "dreamlike") ride the tail; they set the grade, not the subject.
  • Keep spoken lines short enough for the clip — roughly a sentence or two per 8 seconds. Bind each line to the action that motivates it: describe the beat, then give the line.
  • Continuity across shots: repeat the character/wardrobe/light description and keep camera language consistent.

Del manual propio de LUVI para esta familia: las mismas reglas que aplican el reescritor del espacio de trabajo y el motor de prompts de MCP.

Otros modelos de esta familia