Kling·by KuaishouImage to video

Kling Video O3 Pro Image-to-Video

Kling Omni Video O3 Image-to-Video transforms static images into dynamic cinematic videos using MVL technology. Professional quality with first/last frame control and audio generation.

Open in workspacekwaivgi/kling-video-o3-pro/image-to-video

Parameters

ParameterTypeDefaultRange or options
duration
duration

The duration of the generated media in seconds (3-15).

select53, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15
multi shot
multi_shot

Whether to enable multi-shot generation.

selectfalsetrue, false
shot type
shot_type

Multi-shot mode. customize = caller provides per-shot prompts; intelligence = model auto-splits the top-level prompt into shots. Required when multi_shot=true.

selectcustomizecustomize, intelligence
sound
sound

Whether to automatically add audio to the generated video.

selecttruetrue, false

Inputs

Files this model takes in addition to the prompt.

InputAcceptsMax files
First Frame
image
image1Required
Last Frame
end_image
image1Optional

Prompting guide for Kling

TARGET MODEL: Kling 3 (Kuaishou video — v3.0 Pro cinematic / Omni O3 native lip-synced audio). It reads CINEMATIC INTENT, not inventories: direct a scene for a cinematographer, never list objects.

Rules that must hold

  • Every shot is the formula, left-to-right: Subject + Subject Movement + Scene + (Camera Language + Lighting + Atmosphere). Who, doing what, where, then how it's shot and lit. A muddy clip is a missing slot, not a bad seed.
  • Concise beats maximal: 60-100 words, 2-5 sentences per shot, one job per sentence. Never an adjective wall. HARD CAPS: the whole prompt at most 2,500 characters (overrun fails the submission); each shot card ~512 characters — if it won't fit in 512, it's two shots.
  • Kill garnish: delete "cinematic, 4K, dramatic, trending". At most 1-2 earned mood words as a tail AFTER the direction is written (serene, intimate, epic).
  • ONE clean camera move per shot, named by its grip term — "slow push-in", "pull-back", "smooth 180° orbit", "handheld with subtle shake", "tracking shot". "Push in then pan right" never works: two moves means two shots. Name no move and the shot drifts — pin it.
  • Camera pace rides the camera verb: lead camera moves with slow / smooth / steady. Action intensity rides the adjective on the ACTION verb, matched to the user's intent ("slow deliberate wave" vs "frantic rapid wave"). Kill dead verbs ("moves", "goes") — load the verb with speed + direction + posture + contact point ("races down a rain-soaked street, weaving between cars").
  • ONE action per shot — if the user's prompt carries more beats, split into a multi-shot board rather than stacking actions in one shot. Stage the action in chronological beats with an ENDPOINT — preparation, main action, follow-through, landed ("sets it down. Soft clink when cup meets saucer."). Open-ended action drifts.
  • Name the light fixture and its direction, never the feeling: "neon signs", "candlelight", "rim light from top-right", "soft key 45° camera-left" — never "good/nice/dramatic lighting".
  • Exclusions: leave unwanted content OUT of the prompt entirely. Never write "no X" in the prompt body — unwanted things are simply not described.

What works best

  • Anchor motion in physics and a described environment — weight, contact, cloth response. Motion in a void distorts; motion that touches a real world stays coherent.
  • Time + weather are lighting presets: "golden hour", "night", "morning mist", "full moon". Atmosphere = particles that make light visible ("dust in the air", "haze", "drifting smoke") with a hard source raked through.
  • Grade in words at the tail: temperature + contrast + format — "warm tones, high contrast, 35mm film grain".
  • Style is a named medium as a tail clause, never "artistic": "Studio Ghibli animation", "film noir lighting and shadows", "1980s VHS camcorder aesthetic", "photorealistic rendering", "cel-shaded anime".
  • Ambient sound lives in the scene slot ("rain tapping on the roof") — audio generates in the same pass; contact words double as SFX.
  • Dialogue (native audio): action FIRST, then the line — stage the physical beat so the model knows whose mouth to move: "The agent slams his hand on the table. [Black-suited Agent, angrily shouting]: \"Where is the truth?\"". Label speakers uniquely and re-cite the exact label every time (a synonym reads as a new person); tag delivery inside the bracket; sequence turns with "Immediately," and hold beats with "Pause." / "Silence."; lay ambience first ("[Sound: low electronic hum and falling rain.]"); match word count to clip length — a 5s shot holds one or two short lines.
  • Multi-shot boards: one clean beat per shot card, each line ordered framing → subject → subject's motion → camera move; connect cuts with "Continuing…"/"then" (match cut), "Immediately" (hard cut), "Camera cuts to…" (explicit); repeat load-bearing subject descriptors word-for-word across cuts ("man in red suit"); roughly 6 shots ceiling, ~512 characters per shot, whole board within the 2,500-character cap.
  • Character consistency: one unique unchanging label per character, cited every beat — kill "he"/"the man"/synonyms; keep the identity block byte-identical when it recurs.

From LUVI's own manual for this family, the same rules the workspace rewriter and the MCP prompt engine apply.

Other models in this family