Gemini Omni·von GoogleVideo edit
Gemini Omni 1.1 Flash Video Edit
Edit an existing video (up to 10 seconds, MP4) from a text instruction: add, remove, replace, or restyle elements with native audio, while everything the prompt does not mention is preserved. Optional 1-10 reference images can introduce a specific subject or style. Output duration and resolution follow the source clip.
google/gemini-omni-1.1-flash/video-editParameter
| Parameter | Typ | Standard | Bereich oder Optionen |
|---|---|---|---|
seed seedThe random seed to use for the generation. -1 means a random seed will be used. | number | -1 | — |
Thinking level thinking_levelHow much internal reasoning the model does before generating. Higher may improve quality on complex prompts but takes longer. Does not affect price. | select | default | default, high, low |
Eingaben
Dateien, die dieses Modell zusätzlich zum Prompt entgegennimmt.
| Eingabe | Akzeptiert | Max. Dateien | |
|---|---|---|---|
Source video video | video | 1 | Erforderlich |
Reference images reference_images | image | 10 | Optional |
Prompt-Leitfaden für Gemini Omni
TARGET MODEL: Gemini Omni Flash (Google — caption-conditioned joint audio-video). Prose captions, not tag salad. There is no negative-prompt field, no system instruction, no weighting syntax, no camera parameter, no motion scalar and no usable seed — the prompt body is the entire control surface, and it should read like a ~45-word caption: "Continuous, unbroken handheld shot of a fluffy tabby cat sitting on a sunny windowsill, looking out into a leafy garden. The cat's tail twitches slowly, and its ears rotate slightly toward ambient noises. Sunbeams illuminate dust motes in the air. Sound design: Gentle breeze, distant bird chirps. No dialogue."
Regeln, die gelten müssen
- Omni CUTS BY DEFAULT. Every single-take prompt carries an explicit continuity clause — "In a single unbroken scene", "One continuous shot, no jump cuts", "oner" — in the FIRST sentence, and again at the end when the shot is compound. Without it the model solves a camera move by cutting, and the soundtrack gets cut with the picture.
- Prescribe the frame, delegate the world: state the shot framing and motion, style, lighting, location and action; do not itemise period detail, wardrobe logic or physics — the model supplies them.
- ONE narrative beat per generation. Timecoded beats ("[0-3s] …", "After 3 seconds, …", "At 0:05 she touches the mirror") choreograph movement inside one continuous take; separate scenes are separate generations. Never chain A-then-B-then-C events into one short clip.
- Exclusions are prose, placed LAST as bare fragments: "No dialogue." · "No extra sound effects." · "No embellishments." · "No music, just realistic real world sound." · "No text overlay on screen." · "Do not show X." A visible object is removed by inverting to a target state — "Make the phone invisible." — never "remove and inpaint", and never a bare-noun list ("wall, frame"): that Veo habit summons the thing here.
- Audio is generated in the same pass, on by default, with no flag: isolate it in its own sentence behind the label "Sound design: …" — sound effects, ambient noise, dialogue. MUSIC MUST BE ASKED FOR, and by source and processing character, not genre ("a low tinny radio broadcast in the background, playing a song").
- Dialogue uses a colon and NO quotation marks: "In a voice that is crisp and clear, with a thoughtful tone and a standard American accent, Clara says: It has to be here." Quotation marks burn the words into the frame as on-screen text — reserve quotes for text that is meant to be visible: 'There is a street sign that says: "OPEN"'. Never carry Veo's quoted-dialogue convention over.
- Write the prompt in English even when the spoken line or on-screen string is not; only English is evaluated.
- Camera is prose from a documented vocabulary: continuity words (one continuous shot · oner · static · locked off · push in · punch in · dolly zoom · natural smartphone zoom · film camera · webcam style), angles (eye-level, low-angle, high-angle, bird's-eye, worm's-eye, Dutch, close-up, extreme close-up, medium, wide/establishing, over-the-shoulder, POV), moves (pan, tilt, dolly in/out, truck, pedestal, zoom, crane, aerial, handheld, whip pan, arc — dolly ≠ zoom), optics only qualitatively (wide-angle, telephoto, shallow depth of field, bokeh, lens flare, rack focus, fisheye), lighting as NAMED LOOKS only (rembrandt, film noir, high-key, low-key, volumetric, backlit silhouette, golden hour, overcast, moonlight, candlelight, harsh fluorescent, neon). No focal lengths, f-stops, shutter values, kelvin or lighting ratios — they are not read.
- No invented syntax: no (term:1.3), no --flags, no BREAK, no <AUDIO_REF>, no motion-strength number. If a control is not in this manual it does not exist.
- Voice editing, ADR and audio input do not exist; never promise lip-sync — the record is empty.
- PHOTOREALISTIC MINORS ARE OFF-LIMITS: Omni refuses them outright ("The model is currently unable to process photorealistic videos of minors") or, when a child prompt slips through, silently renders the child as an adult (measured 2026-09-09: a 5-year-old came back at 1.08× her mother's height). Never describe a photoreal child, baby or teen. When the story needs one, say so and offer a stylised look that carries the character — animation, claymation, illustration — or keep the child out of frame (hands, a voice, a silhouette at a distance).
Was am besten funktioniert
- The five DeepMind slots in this order: shot framing and motion · style · lighting · location · action — then subject micro-motion (tail twitches, ears rotate, dust motes) and light behaviour, then the Sound design sentence, then terminal negations.
- Detail buys control in GENERATION ("the more detail you add, the more control you'll have"); brevity buys control in EDITING — the two doctrines invert, so never pad an edit instruction.
- Compound and sequential camera moves are fine inside one beat ("dollies forward while simultaneously zooming out", "close-up on his shoes, quickly tilting up to medium shot, then widening") PROVIDED the continuity clause closes the sentence.
- For on-screen text: name the surface, quote the literal string, state duration and treatment; one word or one surface at a time ('One word on the screen at a time: "did, you, know…". Each word appears for 1s.'). Typeface, weight and position are not promptable — clean plates ordered with "No text overlay on screen." belong to post.
- Cross-clip identity has no seed: repeat the entire unchanged character block, including the voice description, verbatim in every prompt of a series.
- Named weaknesses to write around: complex fast motion (reduce to one beat, move the camera instead of the subject), burned-in type, consistency across a long edit chain.
Aus LUVIs eigenem Handbuch für diese Familie; dieselben Regeln wendet der Rewriter im Arbeitsbereich und die MCP-Prompt-Engine an.