Gemini Omni·by GoogleText to video

Gemini Omni 1.1 Flash Text-to-Video

Generate cinematic video from text with Gemini Omni 1.1 Flash, Google DeepMind's natively multimodal model. Every clip is rendered with synchronized native audio (dialogue, ambience, music). 3-10 seconds; resolution from a fast 360p draft up to 4K; prompts up to 20,000 characters.

Open in workspacegoogle/gemini-omni-1.1-flash/text-to-video

Parameters

ParameterTypeDefaultRange or options
duration
duration

The duration of the generated video in seconds. Billed per second.

slider103 – 10
aspect ratio
aspect_ratio

The aspect ratio of the generated video.

select16:916:9, 9:16
Resolution
resolution

360p is a fast, low-cost draft for previewing a shot; 1080p and 4k are upscaled from the natively generated frames. Price scales with resolution.

select720p360p, 720p, 1080p, 4k
seed
seed

The random seed to use for the generation. -1 means a random seed will be used.

number-1
Thinking level
thinking_level

How much internal reasoning the model does before generating. Higher may improve quality on complex prompts but takes longer. Does not affect price.

selectdefaultdefault, high, low

Inputs

No reference files: this model works from the prompt alone.

Prompting guide for Gemini Omni

TARGET MODEL: Gemini Omni Flash (Google — caption-conditioned joint audio-video). Prose captions, not tag salad. There is no negative-prompt field, no system instruction, no weighting syntax, no camera parameter, no motion scalar and no usable seed — the prompt body is the entire control surface, and it should read like a ~45-word caption: "Continuous, unbroken handheld shot of a fluffy tabby cat sitting on a sunny windowsill, looking out into a leafy garden. The cat's tail twitches slowly, and its ears rotate slightly toward ambient noises. Sunbeams illuminate dust motes in the air. Sound design: Gentle breeze, distant bird chirps. No dialogue."

Rules that must hold

  • Omni CUTS BY DEFAULT. Every single-take prompt carries an explicit continuity clause — "In a single unbroken scene", "One continuous shot, no jump cuts", "oner" — in the FIRST sentence, and again at the end when the shot is compound. Without it the model solves a camera move by cutting, and the soundtrack gets cut with the picture.
  • Prescribe the frame, delegate the world: state the shot framing and motion, style, lighting, location and action; do not itemise period detail, wardrobe logic or physics — the model supplies them.
  • ONE narrative beat per generation. Timecoded beats ("[0-3s] …", "After 3 seconds, …", "At 0:05 she touches the mirror") choreograph movement inside one continuous take; separate scenes are separate generations. Never chain A-then-B-then-C events into one short clip.
  • Exclusions are prose, placed LAST as bare fragments: "No dialogue." · "No extra sound effects." · "No embellishments." · "No music, just realistic real world sound." · "No text overlay on screen." · "Do not show X." A visible object is removed by inverting to a target state — "Make the phone invisible." — never "remove and inpaint", and never a bare-noun list ("wall, frame"): that Veo habit summons the thing here.
  • Audio is generated in the same pass, on by default, with no flag: isolate it in its own sentence behind the label "Sound design: …" — sound effects, ambient noise, dialogue. MUSIC MUST BE ASKED FOR, and by source and processing character, not genre ("a low tinny radio broadcast in the background, playing a song").
  • Dialogue uses a colon and NO quotation marks: "In a voice that is crisp and clear, with a thoughtful tone and a standard American accent, Clara says: It has to be here." Quotation marks burn the words into the frame as on-screen text — reserve quotes for text that is meant to be visible: 'There is a street sign that says: "OPEN"'. Never carry Veo's quoted-dialogue convention over.
  • Write the prompt in English even when the spoken line or on-screen string is not; only English is evaluated.
  • Camera is prose from a documented vocabulary: continuity words (one continuous shot · oner · static · locked off · push in · punch in · dolly zoom · natural smartphone zoom · film camera · webcam style), angles (eye-level, low-angle, high-angle, bird's-eye, worm's-eye, Dutch, close-up, extreme close-up, medium, wide/establishing, over-the-shoulder, POV), moves (pan, tilt, dolly in/out, truck, pedestal, zoom, crane, aerial, handheld, whip pan, arc — dolly ≠ zoom), optics only qualitatively (wide-angle, telephoto, shallow depth of field, bokeh, lens flare, rack focus, fisheye), lighting as NAMED LOOKS only (rembrandt, film noir, high-key, low-key, volumetric, backlit silhouette, golden hour, overcast, moonlight, candlelight, harsh fluorescent, neon). No focal lengths, f-stops, shutter values, kelvin or lighting ratios — they are not read.
  • No invented syntax: no (term:1.3), no --flags, no BREAK, no <AUDIO_REF>, no motion-strength number. If a control is not in this manual it does not exist.
  • Voice editing, ADR and audio input do not exist; never promise lip-sync — the record is empty.
  • PHOTOREALISTIC MINORS ARE OFF-LIMITS: Omni refuses them outright ("The model is currently unable to process photorealistic videos of minors") or, when a child prompt slips through, silently renders the child as an adult (measured 2026-09-09: a 5-year-old came back at 1.08× her mother's height). Never describe a photoreal child, baby or teen. When the story needs one, say so and offer a stylised look that carries the character — animation, claymation, illustration — or keep the child out of frame (hands, a voice, a silhouette at a distance).

What works best

  • The five DeepMind slots in this order: shot framing and motion · style · lighting · location · action — then subject micro-motion (tail twitches, ears rotate, dust motes) and light behaviour, then the Sound design sentence, then terminal negations.
  • Detail buys control in GENERATION ("the more detail you add, the more control you'll have"); brevity buys control in EDITING — the two doctrines invert, so never pad an edit instruction.
  • Compound and sequential camera moves are fine inside one beat ("dollies forward while simultaneously zooming out", "close-up on his shoes, quickly tilting up to medium shot, then widening") PROVIDED the continuity clause closes the sentence.
  • For on-screen text: name the surface, quote the literal string, state duration and treatment; one word or one surface at a time ('One word on the screen at a time: "did, you, know…". Each word appears for 1s.'). Typeface, weight and position are not promptable — clean plates ordered with "No text overlay on screen." belong to post.
  • Cross-clip identity has no seed: repeat the entire unchanged character block, including the voice description, verbatim in every prompt of a series.
  • Named weaknesses to write around: complex fast motion (reduce to one beat, move the camera instead of the subject), burned-in type, consistency across a long edit chain.

From LUVI's own manual for this family, the same rules the workspace rewriter and the MCP prompt engine apply.

Other models in this family