Gemini Omni·de GoogleReference to video

Gemini Omni Flash Reference-to-Video

Generate a new scene that carries a referenced character, object, or style from 1-10 reference images. The best variant for keeping a recurring character or a consistent brand look across a series. 3-10 seconds, 720p.

Abrir en el espacio de trabajogoogle/gemini-omni-flash/reference-to-video

Parámetros

ParámetroTipoPor defectoRango u opciones
duration
duration

The duration of the generated video in seconds. Billed per second.

slider103 – 10
aspect ratio
aspect_ratio

The aspect ratio of the generated video.

select16:916:9, 9:16
resolution
resolution

The resolution of the generated video.

select720p720p
seed
seed

The random seed to use for the generation. -1 means a random seed will be used.

number-1
Thinking level
thinking_level

How much internal reasoning the model does before generating. Higher may improve quality on complex prompts but takes longer. Does not affect price.

selectdefaultdefault, high, low

Entradas

Archivos que este modelo recibe además del prompt.

EntradaAceptaMáx. de archivos
images
images
image10Obligatoria

Guía de prompts para Gemini Omni

TARGET MODEL: Gemini Omni Flash (Google — caption-conditioned joint audio-video). Prose captions, not tag salad. There is no negative-prompt field, no system instruction, no weighting syntax, no camera parameter, no motion scalar and no usable seed — the prompt body is the entire control surface, and it should read like a ~45-word caption: "Continuous, unbroken handheld shot of a fluffy tabby cat sitting on a sunny windowsill, looking out into a leafy garden. The cat's tail twitches slowly, and its ears rotate slightly toward ambient noises. Sunbeams illuminate dust motes in the air. Sound design: Gentle breeze, distant bird chirps. No dialogue."

Reglas que deben cumplirse

  • Omni CUTS BY DEFAULT. Every single-take prompt carries an explicit continuity clause — "In a single unbroken scene", "One continuous shot, no jump cuts", "oner" — in the FIRST sentence, and again at the end when the shot is compound. Without it the model solves a camera move by cutting, and the soundtrack gets cut with the picture.
  • Prescribe the frame, delegate the world: state the shot framing and motion, style, lighting, location and action; do not itemise period detail, wardrobe logic or physics — the model supplies them.
  • ONE narrative beat per generation. Timecoded beats ("[0-3s] …", "After 3 seconds, …", "At 0:05 she touches the mirror") choreograph movement inside one continuous take; separate scenes are separate generations. Never chain A-then-B-then-C events into one short clip.
  • Exclusions are prose, placed LAST as bare fragments: "No dialogue." · "No extra sound effects." · "No embellishments." · "No music, just realistic real world sound." · "No text overlay on screen." · "Do not show X." A visible object is removed by inverting to a target state — "Make the phone invisible." — never "remove and inpaint", and never a bare-noun list ("wall, frame"): that Veo habit summons the thing here.
  • Audio is generated in the same pass, on by default, with no flag: isolate it in its own sentence behind the label "Sound design: …" — sound effects, ambient noise, dialogue. MUSIC MUST BE ASKED FOR, and by source and processing character, not genre ("a low tinny radio broadcast in the background, playing a song").
  • Dialogue uses a colon and NO quotation marks: "In a voice that is crisp and clear, with a thoughtful tone and a standard American accent, Clara says: It has to be here." Quotation marks burn the words into the frame as on-screen text — reserve quotes for text that is meant to be visible: 'There is a street sign that says: "OPEN"'. Never carry Veo's quoted-dialogue convention over.
  • Write the prompt in English even when the spoken line or on-screen string is not; only English is evaluated.
  • Camera is prose from a documented vocabulary: continuity words (one continuous shot · oner · static · locked off · push in · punch in · dolly zoom · natural smartphone zoom · film camera · webcam style), angles (eye-level, low-angle, high-angle, bird's-eye, worm's-eye, Dutch, close-up, extreme close-up, medium, wide/establishing, over-the-shoulder, POV), moves (pan, tilt, dolly in/out, truck, pedestal, zoom, crane, aerial, handheld, whip pan, arc — dolly ≠ zoom), optics only qualitatively (wide-angle, telephoto, shallow depth of field, bokeh, lens flare, rack focus, fisheye), lighting as NAMED LOOKS only (rembrandt, film noir, high-key, low-key, volumetric, backlit silhouette, golden hour, overcast, moonlight, candlelight, harsh fluorescent, neon). No focal lengths, f-stops, shutter values, kelvin or lighting ratios — they are not read.
  • No invented syntax: no (term:1.3), no --flags, no BREAK, no <AUDIO_REF>, no motion-strength number. If a control is not in this manual it does not exist.
  • Voice editing, ADR and audio input do not exist; never promise lip-sync — the record is empty.
  • PHOTOREALISTIC MINORS ARE OFF-LIMITS: Omni refuses them outright ("The model is currently unable to process photorealistic videos of minors") or, when a child prompt slips through, silently renders the child as an adult (measured 2026-09-09: a 5-year-old came back at 1.08× her mother's height). Never describe a photoreal child, baby or teen. When the story needs one, say so and offer a stylised look that carries the character — animation, claymation, illustration — or keep the child out of frame (hands, a voice, a silhouette at a distance).

Lo que mejor funciona

  • The five DeepMind slots in this order: shot framing and motion · style · lighting · location · action — then subject micro-motion (tail twitches, ears rotate, dust motes) and light behaviour, then the Sound design sentence, then terminal negations.
  • Detail buys control in GENERATION ("the more detail you add, the more control you'll have"); brevity buys control in EDITING — the two doctrines invert, so never pad an edit instruction.
  • Compound and sequential camera moves are fine inside one beat ("dollies forward while simultaneously zooming out", "close-up on his shoes, quickly tilting up to medium shot, then widening") PROVIDED the continuity clause closes the sentence.
  • For on-screen text: name the surface, quote the literal string, state duration and treatment; one word or one surface at a time ('One word on the screen at a time: "did, you, know…". Each word appears for 1s.'). Typeface, weight and position are not promptable — clean plates ordered with "No text overlay on screen." belong to post.
  • Cross-clip identity has no seed: repeat the entire unchanged character block, including the voice description, verbatim in every prompt of a series.
  • Named weaknesses to write around: complex fast motion (reduce to one beat, move the camera instead of the subject), burned-in type, consistency across a long edit chain.

Del manual propio de LUVI para esta familia: las mismas reglas que aplican el reescritor del espacio de trabajo y el motor de prompts de MCP.

Otros modelos de esta familia