Kling·by KuaishouImage to video
Kling v3.0 Pro Image-to-Video
Kling v3.0 Professional Image-to-Video model by Kuaishou. Premium quality video generation from images with advanced features.
kwaivgi/kling-v3.0-pro/image-to-videoParameters
| Parameter | Type | Default | Range or options |
|---|---|---|---|
duration durationThe duration of the generated media in seconds (3-15). | select | 5 | 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 |
resolution resolutionThe resolution of the generated video. | textarea | 1080P | — |
cfg scale cfg_scaleFlexibility in video generation; The higher the value, the lower the model's degree of flexibility, and the stronger the relevance to the user's prompt. | slider | 0.5 | 0 – 1 |
multi shot multi_shotWhether to enable multi-shot generation. | select | false | true, false |
shot type shot_typeMulti-shot mode. customize = caller provides per-shot prompts; intelligence = model auto-splits the top-level prompt into shots. Required when multi_shot=true. | select | customize | customize, intelligence |
sound soundWhether sound is generated simultaneously when generating a video. | select | true | true, false |
Inputs
Files this model takes in addition to the prompt.
| Input | Accepts | Max files | |
|---|---|---|---|
First Frame image | image | 1 | Required |
Last Frame end_image | image | 1 | Optional |
Prompting guide for Kling
TARGET MODEL: Kling 3 (Kuaishou video — v3.0 Pro cinematic / Omni O3 native lip-synced audio). It reads CINEMATIC INTENT, not inventories: direct a scene for a cinematographer, never list objects.
Rules that must hold
- Every shot is the formula, left-to-right: Subject + Subject Movement + Scene + (Camera Language + Lighting + Atmosphere). Who, doing what, where, then how it's shot and lit. A muddy clip is a missing slot, not a bad seed.
- Concise beats maximal: 60-100 words, 2-5 sentences per shot, one job per sentence. Never an adjective wall. HARD CAPS: the whole prompt at most 2,500 characters (overrun fails the submission); each shot card ~512 characters — if it won't fit in 512, it's two shots.
- Kill garnish: delete "cinematic, 4K, dramatic, trending". At most 1-2 earned mood words as a tail AFTER the direction is written (serene, intimate, epic).
- ONE clean camera move per shot, named by its grip term — "slow push-in", "pull-back", "smooth 180° orbit", "handheld with subtle shake", "tracking shot". "Push in then pan right" never works: two moves means two shots. Name no move and the shot drifts — pin it.
- Camera pace rides the camera verb: lead camera moves with slow / smooth / steady. Action intensity rides the adjective on the ACTION verb, matched to the user's intent ("slow deliberate wave" vs "frantic rapid wave"). Kill dead verbs ("moves", "goes") — load the verb with speed + direction + posture + contact point ("races down a rain-soaked street, weaving between cars").
- ONE action per shot — if the user's prompt carries more beats, split into a multi-shot board rather than stacking actions in one shot. Stage the action in chronological beats with an ENDPOINT — preparation, main action, follow-through, landed ("sets it down. Soft clink when cup meets saucer."). Open-ended action drifts.
- Name the light fixture and its direction, never the feeling: "neon signs", "candlelight", "rim light from top-right", "soft key 45° camera-left" — never "good/nice/dramatic lighting".
- Exclusions: leave unwanted content OUT of the prompt entirely. Never write "no X" in the prompt body — unwanted things are simply not described.
What works best
- Anchor motion in physics and a described environment — weight, contact, cloth response. Motion in a void distorts; motion that touches a real world stays coherent.
- Time + weather are lighting presets: "golden hour", "night", "morning mist", "full moon". Atmosphere = particles that make light visible ("dust in the air", "haze", "drifting smoke") with a hard source raked through.
- Grade in words at the tail: temperature + contrast + format — "warm tones, high contrast, 35mm film grain".
- Style is a named medium as a tail clause, never "artistic": "Studio Ghibli animation", "film noir lighting and shadows", "1980s VHS camcorder aesthetic", "photorealistic rendering", "cel-shaded anime".
- Ambient sound lives in the scene slot ("rain tapping on the roof") — audio generates in the same pass; contact words double as SFX.
- Dialogue (native audio): action FIRST, then the line — stage the physical beat so the model knows whose mouth to move: "The agent slams his hand on the table. [Black-suited Agent, angrily shouting]: \"Where is the truth?\"". Label speakers uniquely and re-cite the exact label every time (a synonym reads as a new person); tag delivery inside the bracket; sequence turns with "Immediately," and hold beats with "Pause." / "Silence."; lay ambience first ("[Sound: low electronic hum and falling rain.]"); match word count to clip length — a 5s shot holds one or two short lines.
- Multi-shot boards: one clean beat per shot card, each line ordered framing → subject → subject's motion → camera move; connect cuts with "Continuing…"/"then" (match cut), "Immediately" (hard cut), "Camera cuts to…" (explicit); repeat load-bearing subject descriptors word-for-word across cuts ("man in red suit"); roughly 6 shots ceiling, ~512 characters per shot, whole board within the 2,500-character cap.
- Character consistency: one unique unchanging label per character, cited every beat — kill "he"/"the man"/synonyms; keep the identity block byte-identical when it recurs.
From LUVI's own manual for this family, the same rules the workspace rewriter and the MCP prompt engine apply.