by GoogleVideo

Gemini Omni

9 models · from 2,200 credits

Models in the Gemini Omni family

ModelModePriceAspect ratios
Gemini Omni 1.1 Flash Image-to-VideoImage to videofrom 2,200 credits16:9 · 9:16Open
Gemini Omni 1.1 Flash Reference-to-VideoReference to videofrom 2,200 credits16:9 · 9:16Open
Gemini Omni 1.1 Flash Text-to-VideoText to videofrom 2,200 credits16:9 · 9:16Open
Gemini Omni 1.1 Flash Video EditVideo editfrom 6,374 creditsOpen
Gemini Omni 1.1 Flash Video ExtendVideo extendfrom 2,374 creditsOpen
Gemini Omni Flash Image-to-VideoImage to videofrom 2,600 credits16:9 · 9:16Open
Gemini Omni Flash Reference-to-VideoReference to videofrom 2,700 credits16:9 · 9:16Open
Gemini Omni Flash Text-to-VideoText to videofrom 2,500 credits16:9 · 9:16Open
Gemini Omni Flash Video EditVideo editfrom 8,400 creditsOpen

Prompting guide for Gemini Omni

TARGET MODEL: Gemini Omni Flash (Google — caption-conditioned joint audio-video). Prose captions, not tag salad. There is no negative-prompt field, no system instruction, no weighting syntax, no camera parameter, no motion scalar and no usable seed — the prompt body is the entire control surface, and it should read like a ~45-word caption: "Continuous, unbroken handheld shot of a fluffy tabby cat sitting on a sunny windowsill, looking out into a leafy garden. The cat's tail twitches slowly, and its ears rotate slightly toward ambient noises. Sunbeams illuminate dust motes in the air. Sound design: Gentle breeze, distant bird chirps. No dialogue."

Rules that must hold

  • Omni CUTS BY DEFAULT. Every single-take prompt carries an explicit continuity clause — "In a single unbroken scene", "One continuous shot, no jump cuts", "oner" — in the FIRST sentence, and again at the end when the shot is compound. Without it the model solves a camera move by cutting, and the soundtrack gets cut with the picture.
  • Prescribe the frame, delegate the world: state the shot framing and motion, style, lighting, location and action; do not itemise period detail, wardrobe logic or physics — the model supplies them.
  • ONE narrative beat per generation. Timecoded beats ("[0-3s] …", "After 3 seconds, …", "At 0:05 she touches the mirror") choreograph movement inside one continuous take; separate scenes are separate generations. Never chain A-then-B-then-C events into one short clip.
  • Exclusions are prose, placed LAST as bare fragments: "No dialogue." · "No extra sound effects." · "No embellishments." · "No music, just realistic real world sound." · "No text overlay on screen." · "Do not show X." A visible object is removed by inverting to a target state — "Make the phone invisible." — never "remove and inpaint", and never a bare-noun list ("wall, frame"): that Veo habit summons the thing here.
  • Audio is generated in the same pass, on by default, with no flag: isolate it in its own sentence behind the label "Sound design: …" — sound effects, ambient noise, dialogue. MUSIC MUST BE ASKED FOR, and by source and processing character, not genre ("a low tinny radio broadcast in the background, playing a song").
  • Dialogue uses a colon and NO quotation marks: "In a voice that is crisp and clear, with a thoughtful tone and a standard American accent, Clara says: It has to be here." Quotation marks burn the words into the frame as on-screen text — reserve quotes for text that is meant to be visible: 'There is a street sign that says: "OPEN"'. Never carry Veo's quoted-dialogue convention over.
  • Write the prompt in English even when the spoken line or on-screen string is not; only English is evaluated.
  • Camera is prose from a documented vocabulary: continuity words (one continuous shot · oner · static · locked off · push in · punch in · dolly zoom · natural smartphone zoom · film camera · webcam style), angles (eye-level, low-angle, high-angle, bird's-eye, worm's-eye, Dutch, close-up, extreme close-up, medium, wide/establishing, over-the-shoulder, POV), moves (pan, tilt, dolly in/out, truck, pedestal, zoom, crane, aerial, handheld, whip pan, arc — dolly ≠ zoom), optics only qualitatively (wide-angle, telephoto, shallow depth of field, bokeh, lens flare, rack focus, fisheye), lighting as NAMED LOOKS only (rembrandt, film noir, high-key, low-key, volumetric, backlit silhouette, golden hour, overcast, moonlight, candlelight, harsh fluorescent, neon). No focal lengths, f-stops, shutter values, kelvin or lighting ratios — they are not read.
  • No invented syntax: no (term:1.3), no --flags, no BREAK, no <AUDIO_REF>, no motion-strength number. If a control is not in this manual it does not exist.
  • Voice editing, ADR and audio input do not exist; never promise lip-sync — the record is empty.
  • PHOTOREALISTIC MINORS ARE OFF-LIMITS: Omni refuses them outright ("The model is currently unable to process photorealistic videos of minors") or, when a child prompt slips through, silently renders the child as an adult (measured 2026-09-09: a 5-year-old came back at 1.08× her mother's height). Never describe a photoreal child, baby or teen. When the story needs one, say so and offer a stylised look that carries the character — animation, claymation, illustration — or keep the child out of frame (hands, a voice, a silhouette at a distance).

What works best

  • The five DeepMind slots in this order: shot framing and motion · style · lighting · location · action — then subject micro-motion (tail twitches, ears rotate, dust motes) and light behaviour, then the Sound design sentence, then terminal negations.
  • Detail buys control in GENERATION ("the more detail you add, the more control you'll have"); brevity buys control in EDITING — the two doctrines invert, so never pad an edit instruction.
  • Compound and sequential camera moves are fine inside one beat ("dollies forward while simultaneously zooming out", "close-up on his shoes, quickly tilting up to medium shot, then widening") PROVIDED the continuity clause closes the sentence.
  • For on-screen text: name the surface, quote the literal string, state duration and treatment; one word or one surface at a time ('One word on the screen at a time: "did, you, know…". Each word appears for 1s.'). Typeface, weight and position are not promptable — clean plates ordered with "No text overlay on screen." belong to post.
  • Cross-clip identity has no seed: repeat the entire unchanged character block, including the voice description, verbatim in every prompt of a series.
  • Named weaknesses to write around: complex fast motion (reduce to one beat, move the camera instead of the subject), burned-in type, consistency across a long edit chain.

From LUVI's own manual for this family, the same rules the workspace rewriter and the MCP prompt engine apply.