Vidu Q3·by ViduImage to video
Vidu Q3-Pro Image-to-video
Vidu Q3-Pro Image-to-Video is an advanced AI video generation model that brings static images to life. Upload a reference image and describe the motion you want — the model generates high-quality video with smooth animation, optional audio, and cinematic quality up to 1080p.
vidu/q3-pro/image-to-videoParameters
| Parameter | Type | Default | Range or options |
|---|---|---|---|
duration durationThe duration of the generated media in seconds. | slider | 5 | 1 – 16 |
resolution resolutionThe resolution of the generated media. | select | 720p | 540p, 720p, 1080p |
seed seedThe random seed to use for the generation. -1 means a random seed will be used. | number | — | — |
bgm bgmThe background music for generating the output. | select | true | true, false |
generate audio generate_audioWhether to generate audio. | select | true | true, false |
Inputs
Files this model takes in addition to the prompt.
| Input | Accepts | Max files | |
|---|---|---|---|
image image | image | 1 | Required |
Prompting guide for Vidu Q3
TARGET MODEL: Vidu Q3 (Shengshu — joint audio-video in one pass: lip-synced dialogue, voiceover, sound effects and music generated with the picture; 1–16 s; native camera control and "smart cuts" multi-shot). Plain descriptive prose. No negative prompt, no weighting syntax; the style and motion-strength parameters are documented as ineffective on Q3, so everything about look and energy is said in words.
Rules that must hold
- Describe subject, action, setting, style, camera movement and mood as prose; the official example is one flowing paragraph ("In an ultra-realistic fashion photography style featuring light blue and pale amber tones, an astronaut in a spacesuit walks through the fog. The background consists of enchanting white and golden lights…"). No tag lists, no bracketed commands, no (word:1.3).
- Look and motion energy are words, not parameters: name the visual style in the sentence ("anime", "ultra-realistic fashion photography") and the amount of motion ("barely moves", "sprints") — the style and movement_amplitude parameters do nothing on Q3.
- Sound is generated with the picture whenever audio is on: write it as content — a spoken line in quotes attributed to a visible speaker, named sound effects tied to visible actions, the music's character. Multi-speaker conversation is supported: label each speaker uniquely and give each line its own sentence.
- Dialogue languages that are supported: English, Japanese, Chinese. Another language may render poorly — keep such a line very short or carry the beat without on-camera speech.
- One primary subject with supporting environmental detail per clip; do not name real people, brands or titles.
- Hard cap 5,000 characters; a clip prompt is a paragraph, not a script.
What works best
- State the camera explicitly with film terms — pans, push-ins, tracking shots, "slow aerial drone shot", "static locked-off" — Q3 honors frame-level camera direction and a named move is more consistent than an implied one.
- Multi-shot is native ("smart cuts"): when the story has beats, write them in order, each with its framing and its sound, and let the model place the cuts; a single-take clip says so ("one continuous shot").
- Environmental sound follows environmental detail: "busy Tokyo crosswalk at night" pulls tires, train horns and signal beeps — describe the world and its sounds together ("campfire crackles… sparks drifting… crickets chirping, occasional owl hoot").
- Match content to seconds: a 5-second clip is one beat; use the 16-second ceiling for a short sequence, not for a longer single action.
From LUVI's own manual for this family, the same rules the workspace rewriter and the MCP prompt engine apply.