How to prompt Kling 3 — and why your prompt never reached the model

Kling 3 reads a shot as subject, movement and scene first, then camera, light and atmosphere: 60 to 100 words, one camera move, one action. The whole prompt is hard-capped at 2,500 characters and an overrun fails the submission before anything renders.

Published:

Six translucent cards glide in a row across a near-black frame, each holding a single teal streak of light, while the last one is held against a narrow indigo threshold and warm amber ripples drift past.

Kling 3 is Kuaishou's video model, and on LUVI the Kling family runs it in two shapes: the v3.0 models for cinematic picture, and the Video O3 models for native audio. Clips run 3 to 15 seconds in 16:9, 9:16 or 1:1, and sound is switched on by default. Before you run anything, the Workspace shows the estimate in credits for the settings you picked. What follows is the prompting guide we run inside LUVI, plus the parts of the model contract that decide whether a prompt is accepted at all.

What order does a Kling 3 prompt follow?

A Kling 3 shot is written left to right as subject, subject movement, scene, then camera language, lighting and atmosphere. Our prompting guide states the rule plainly: "A muddy clip is a missing slot, not a bad seed." Keep one shot to 60–100 words across two to five sentences, one job per sentence.

  • Subject and movement first. Load the action verb with speed, direction and contact: "races down a rain-soaked street, weaving between cars" rather than "moves".
  • One camera move, named by its grip term. "Slow push-in", "smooth 180° orbit", "handheld with subtle shake". Our guide is blunt about the alternative: "'Push in then pan right' never works: two moves means two shots."
  • Name the light fixture and its direction, never the feeling. "Rim light from top-right" works; "dramatic lighting" does not.
  • Delete the garnish. "Cinematic, 4K, dramatic, trending" earn nothing. One or two mood words can sit at the tail, after the direction is written.
  • Grade in words at the end: temperature, contrast and format, such as "warm tones, high contrast, 35mm film grain".

That gives a single shot like this one, for Kling v3.0 Pro Text-to-Video:

A cellist in a grey wool coat draws a slow bow across the strings, shoulders easing down with the phrase. She sits alone on the stage of an empty theatre, rows of red velvet seats fading into the dark behind her. Slow push-in from the third row. Hard key from a single overhead work light, rim light from top-right, dust drifting through the beam. Warm tones, high contrast, 35mm film grain.

Why did my Kling 3 prompt fail to submit?

Kling 3 hard-caps the prompt at 2,500 characters, and going over fails the submission rather than trimming the text. Each shot card should stay around 512 characters; if a shot won't fit, it is two shots. This is the one rule the machine enforces for you: run a 2,565-character prompt through check_prompt on the connector and it comes back with pass: false and a single char-cap issue citing the Kling guide — "2565 characters — Kling 3 hard-caps the prompt at 2500; overrun fails. Tighten it."

The craft rules are not enforced that way. We put a deliberately bad prompt — "Cinematic 4K dramatic shot of a man walking down a street, no rain, no cars, push in then pan right, beautiful dramatic lighting, trending, masterpiece, highly detailed, 8k" — through the same checker on 16 September 2026, and it returned pass: true with no transforms, although it breaks four rules the guide states. Treat the checker as a length gate, and the guide as the thing you read.

How do I keep something out of a Kling 3 shot?

Leave it out of the prompt entirely. Kling 3 has no negative prompt field on LUVI, and our guide's instruction is to never write "no X" in the prompt body: unwanted things are simply not described. Write the empty road, not the missing cars. Naming what you don't want puts the word in front of the model, and the checker will not catch it for you.

How do I write dialogue and sound for Kling 3?

Stage the physical beat first, then give the line, so the model knows whose mouth to move. The Video O3 models generate audio in the same pass as the picture, and the guide's pattern is an action, then a bracketed speaker label with the delivery, then the line:

[Sound: rain on a tin roof, the low hum of a chest freezer.] The Night Clerk slides a paper bag across the counter and lets go of it. [Night Clerk, flat and tired]: "You're the third one tonight." Pause. The Night Clerk wipes both hands on a green apron. Locked-off wide from behind the register, one fluorescent tube overhead with a flickering bulb. Cool tones, high contrast.

  • Lay the ambience first, in its own bracket, before anyone speaks.
  • Label each speaker uniquely and repeat the label word for word. A synonym reads as a new person, so "the clerk" and "the man" break the voice you just established.
  • Sequence the turns with "Immediately," and hold the beats with "Pause." or "Silence."
  • Match the words to the length. A five-second shot holds one or two short lines.
  • Ambient sound belongs in the scene slot — "rain tapping on the roof" — because contact words double as sound effects.

Sound is different here from the other family we have written about: on Kling 3, sound defaults to true on both the v3.0 and the Video O3 text-to-video models, while Veo 3.1 keeps audio off until you switch it on. Kuaishou announced 3.0 with "native audio generation across multiple languages, dialects, and accents" in its launch release on 5 February 2026.

How do I get more than one shot out of one generation?

Turn on multi-shot and hand Kling 3 one clean beat per shot card. The model contract is specific, and it matches the guide's advice: multi_shot switches the mode on, shot_type chooses between customize (you write each shot) and intelligence (the model splits your prompt), and multi_prompt carries the storyboard — at most six shots, with each shot's duration adding up to the clip's total duration. The guide's "roughly 6 shots ceiling" is the same six the contract accepts.

  • Order every shot line the same way: framing, subject, the subject's motion, camera move.
  • Connect the cuts in words: "Continuing…" or "then" for a match cut, "Immediately" for a hard cut, "Camera cuts to…" when you want it explicit.
  • Repeat load-bearing descriptors byte for byte across cuts — "man in red suit" stays "man in red suit" — because character identity travels through the exact label, not through pronouns.

Kuaishou describes the same feature from its side: the Video 3.0 Omni model "rolls out a multi-shot storyboard feature that allows users to generate professional shots where they can specify the duration, shot size, perspective, narrative content and camera movements for each shot in storyboarding".

Which Kling 3 model should I pick?

Pick by what you need to control, since the four text-to-video models differ in exactly one parameter and in what the audio is for.

  • Cinematic picture, with prompt adherence you can tune: Kling v3.0 Pro Text-to-Video and Kling v3.0 Std Text-to-Video carry cfg_scale, a 0–1 slider that trades flexibility for closeness to the prompt, at 0.5 by default.
  • Dialogue and native audio: Kling Video O3 Pro Text-to-Video and Kling Video O3 Std Text-to-Video drop cfg_scale and lean on the Omni side of the family, which our guide describes as "native lip-synced audio".
  • Starting from a picture: the same split holds for the image-to-video models, and the O3 models add reference-to-video and video-edit.
  • Every one of them takes 3 to 15 seconds and 16:9, 9:16 or 1:1, with no resolution parameter to set and no seed to pin.

Std and Pro sit at different prices on every mode, and the estimate for the settings you chose appears before the run starts.

How we checked

  • What we read: LUVI's public prompting guide for Kling 3, and the parameter contract of the four Kling 3 text-to-video models through the connector, on 16 September 2026.
  • What we measured: check_prompt on four prompts for kwaivgi/kling-v3.0-pro/text-to-video and kwaivgi/kling-video-o3-pro/text-to-video — the two example prompts above (both pass: true, no transforms), one 2,565-character prompt (pass: false, char-cap), and one garnish-and-negation prompt (pass: true, no transforms).
  • Sound defaults: read from the model schemas of the Kling 3 and Veo 3.1 Fast text-to-video models on the same day.
  • Drawbacks: no outputs were generated for this post, so it makes no claim about output quality, and the character cap is the only rule we saw the checker enforce.

In LUVI

In the Workspace, pick a Kling model, write the shot in the order above, and read the estimate before you run. The feather button next to the prompt rewrites a draft into the Kling dialect for free, and what a prompting guide is explains where these rules come from. Working from Claude or ChatGPT instead, the LUVI connector returns the same guide and the exact parameters, multi_prompt included. To try it, create a LUVI account.

Sources

More from Guides