Veo·par GoogleText to video
Veo 3.1 Lite Text-to-Video
Veo 3.1 Lite: the most economical Veo tier. Text-to-video with strong price/performance for high-volume applications.
google/veo3.1-lite/text-to-videoParamètres
| Paramètre | Type | Par défaut | Plage ou options |
|---|---|---|---|
duration durationLength of the generated video in seconds. Must be 8 when using 1080p resolution. | select | 8 | 4, 6, 8 |
aspect ratio aspect_ratioThe aspect ratio of the generated video. | select | 16:9 | 16:9, 9:16 |
resolution resolutionOutput resolution. 1080p only supports 8s duration and costs more per second. Veo 3.1 Lite does not support 4K. | select | 720p | 720p, 1080p |
seed seedRandom seed. Does not guarantee determinism but may improve repeatability. | number | — | — |
Entrées
Aucun fichier de référence : ce modèle travaille à partir du prompt seul.
Guide de prompt pour Veo
TARGET MODEL: Veo (Google DeepMind video — the shot arrives already scored: picture AND audio generate in the same pass, lip-synced dialogue, SFX, ambience, music). Write a described shot, not tags — a camera, a subject doing something, a world, a look.
Règles à respecter
- Slot order (Google's published, rewarded order): Cinematography → Subject → Action → Context → Style & Ambiance. Lead with the camera, name the subject, give it a verb, place it, then set the stock and the light.
- Front-load the subject and its action — the first clause sets adherence.
- ONE clean beat per clip. A beat is a single action that completes inside the clip. For more than one move, timestamp the beats — "[00:00-00:02] …", "[00:02-00:04] …" — one shot + one sound cue per bracket. Never cram a scene into one unbroken description.
- Camera moves by their PROPER classical names — "dolly shot", "tracking shot", "crane shot", "aerial view", "slow pan", "tilt", "arc", "POV shot"; shot size and angle by name ("wide shot", "close-up", "extreme close-up", "low angle", "two-shot"); optics by name ("shallow depth of field", "wide-angle lens", "macro lens", "deep focus"). One move per beat — stacked moves muddy all of them.
- In-prose exclusions are DESCRIBED ABSENCES, never prohibitions: "a desolate landscape with no buildings or roads" describes emptiness; never write bare "no X" / "don't" commands in the shot.
- Audio is written INTO the shot: the exact spoken line in quotes (it will be lip-synced), attributed to a distinct speaker — 'A woman says, "We have to leave now."'; sound effects cued by name — "SFX: thunder cracks in the distance"; room tone set — "Ambient noise: the quiet hum of a starship bridge"; music named — "soft piano score builds under the scene"; silence is directable — state it and specify ambience only ("no music, only the wind").
Ce qui fonctionne le mieux
- Use DeepMind's seven nameable elements as a checklist so the model isn't guessing defaults: shot & motion, style, lighting, character, location, action, dialogue.
- Resolve the subject to matter (age, wardrobe, material, wear), then give it one clear action that UNFOLDS over the clip — describe how the movement progresses, not a frozen pose.
- Physics is a prompt: name weight, contact, momentum. Vague verbs drift.
- Lighting = source + direction + quality + temperature, plus time-of-day/weather ("golden hour", "overcast", "neon night"). Named setups land as intent.
- One dominant look — genre, era, film-stock reference, animation style — as a clause, never a stack. Ambiance words ("tense", "serene", "dreamlike") ride the tail; they set the grade, not the subject.
- Keep spoken lines short enough for the clip — roughly a sentence or two per 8 seconds. Bind each line to the action that motivates it: describe the beat, then give the line.
- Continuity across shots: repeat the character/wardrobe/light description and keep camera language consistent.
Issu du manuel LUVI pour cette famille : les mêmes règles qu'appliquent le réécriveur de l'espace de travail et le moteur de prompt MCP.