Gemini Omni·by GoogleText to video
Gemini Omni 1.1 Flash Text-to-Video
Generate cinematic video from text with Gemini Omni 1.1 Flash, Google DeepMind's natively multimodal model. Every clip is rendered with synchronized native audio (dialogue, ambience, music). 3-10 seconds; resolution from a fast 360p draft up to 4K; prompts up to 20,000 characters.
google/gemini-omni-1.1-flash/text-to-videoParameters
| Parameter | Type | Default | Range or options |
|---|---|---|---|
duration durationThe duration of the generated video in seconds. Billed per second. | slider | 10 | 3 – 10 |
aspect ratio aspect_ratioThe aspect ratio of the generated video. | select | 16:9 | 16:9, 9:16 |
Resolution resolution360p is a fast, low-cost draft for previewing a shot; 1080p and 4k are upscaled from the natively generated frames. Price scales with resolution. | select | 720p | 360p, 720p, 1080p, 4k |
seed seedThe random seed to use for the generation. -1 means a random seed will be used. | number | -1 | — |
Thinking level thinking_levelHow much internal reasoning the model does before generating. Higher may improve quality on complex prompts but takes longer. Does not affect price. | select | default | default, high, low |
Inputs
No reference files: this model works from the prompt alone.
Prompting guide for Gemini Omni
TARGET MODEL: Gemini Omni Flash (Google — caption-conditioned joint audio-video). Prose captions, not tag salad. There is no negative-prompt field, no system instruction, no weighting syntax, no camera parameter, no motion scalar and no usable seed — the prompt body is the entire control surface, and it should read like a ~45-word caption: "Continuous, unbroken handheld shot of a fluffy tabby cat sitting on a sunny windowsill, looking out into a leafy garden. The cat's tail twitches slowly, and its ears rotate slightly toward ambient noises. Sunbeams illuminate dust motes in the air. Sound design: Gentle breeze, distant bird chirps. No dialogue."
Rules that must hold
- Omni CUTS BY DEFAULT. Every single-take prompt carries an explicit continuity clause — "In a single unbroken scene", "One continuous shot, no jump cuts", "oner" — in the FIRST sentence, and again at the end when the shot is compound. Without it the model solves a camera move by cutting, and the soundtrack gets cut with the picture.
- Prescribe the frame, delegate the world: state the shot framing and motion, style, lighting, location and action; do not itemise period detail, wardrobe logic or physics — the model supplies them.
- ONE narrative beat per generation. Timecoded beats ("[0-3s] …", "After 3 seconds, …", "At 0:05 she touches the mirror") choreograph movement inside one continuous take; separate scenes are separate generations. Never chain A-then-B-then-C events into one short clip.
- Exclusions are prose, placed LAST as bare fragments: "No dialogue." · "No extra sound effects." · "No embellishments." · "No music, just realistic real world sound." · "No text overlay on screen." · "Do not show X." A visible object is removed by inverting to a target state — "Make the phone invisible." — never "remove and inpaint", and never a bare-noun list ("wall, frame"): that Veo habit summons the thing here.
- Audio is generated in the same pass, on by default, with no flag: isolate it in its own sentence behind the label "Sound design: …" — sound effects, ambient noise, dialogue. MUSIC MUST BE ASKED FOR, and by source and processing character, not genre ("a low tinny radio broadcast in the background, playing a song").
- Dialogue uses a colon and NO quotation marks: "In a voice that is crisp and clear, with a thoughtful tone and a standard American accent, Clara says: It has to be here." Quotation marks burn the words into the frame as on-screen text — reserve quotes for text that is meant to be visible: 'There is a street sign that says: "OPEN"'. Never carry Veo's quoted-dialogue convention over.
- Write the prompt in English even when the spoken line or on-screen string is not; only English is evaluated.
- Camera is prose from a documented vocabulary: continuity words (one continuous shot · oner · static · locked off · push in · punch in · dolly zoom · natural smartphone zoom · film camera · webcam style), angles (eye-level, low-angle, high-angle, bird's-eye, worm's-eye, Dutch, close-up, extreme close-up, medium, wide/establishing, over-the-shoulder, POV), moves (pan, tilt, dolly in/out, truck, pedestal, zoom, crane, aerial, handheld, whip pan, arc — dolly ≠ zoom), optics only qualitatively (wide-angle, telephoto, shallow depth of field, bokeh, lens flare, rack focus, fisheye), lighting as NAMED LOOKS only (rembrandt, film noir, high-key, low-key, volumetric, backlit silhouette, golden hour, overcast, moonlight, candlelight, harsh fluorescent, neon). No focal lengths, f-stops, shutter values, kelvin or lighting ratios — they are not read.
- No invented syntax: no (term:1.3), no --flags, no BREAK, no <AUDIO_REF>, no motion-strength number. If a control is not in this manual it does not exist.
- Voice editing, ADR and audio input do not exist; never promise lip-sync — the record is empty.
- PHOTOREALISTIC MINORS ARE OFF-LIMITS: Omni refuses them outright ("The model is currently unable to process photorealistic videos of minors") or, when a child prompt slips through, silently renders the child as an adult (measured 2026-09-09: a 5-year-old came back at 1.08× her mother's height). Never describe a photoreal child, baby or teen. When the story needs one, say so and offer a stylised look that carries the character — animation, claymation, illustration — or keep the child out of frame (hands, a voice, a silhouette at a distance).
What works best
- The five DeepMind slots in this order: shot framing and motion · style · lighting · location · action — then subject micro-motion (tail twitches, ears rotate, dust motes) and light behaviour, then the Sound design sentence, then terminal negations.
- Detail buys control in GENERATION ("the more detail you add, the more control you'll have"); brevity buys control in EDITING — the two doctrines invert, so never pad an edit instruction.
- Compound and sequential camera moves are fine inside one beat ("dollies forward while simultaneously zooming out", "close-up on his shoes, quickly tilting up to medium shot, then widening") PROVIDED the continuity clause closes the sentence.
- For on-screen text: name the surface, quote the literal string, state duration and treatment; one word or one surface at a time ('One word on the screen at a time: "did, you, know…". Each word appears for 1s.'). Typeface, weight and position are not promptable — clean plates ordered with "No text overlay on screen." belong to post.
- Cross-clip identity has no seed: repeat the entire unchanged character block, including the voice description, verbatim in every prompt of a series.
- Named weaknesses to write around: complex fast motion (reduce to one beat, move the camera instead of the subject), burned-in type, consistency across a long edit chain.
From LUVI's own manual for this family, the same rules the workspace rewriter and the MCP prompt engine apply.