Vidu Q3·par ViduReference to video

Vidu Q3-Mix Reference to Video

Vidu Q3-Mix Reference-to-Video generates videos from 1-4 reference images with consistent subjects. Offers strong visual quality with intelligent scene transitions, smooth dynamic effects, and audio support up to 1080p.

Ouvrir dans l'espace de travailvidu/q3-mix/reference-to-video

Paramètres

ParamètreTypePar défautPlage ou options
duration
duration

The duration of the generated video in seconds.

slider51 – 16
aspect ratio
aspect_ratio

The aspect ratio of the output video.

select16:916:9, 9:16, 3:4, 4:3, 1:1
resolution
resolution

The resolution of the generated video.

select720p720p, 1080p
seed
seed

The random seed to use for the generation. Set -1 for random.

number0
generate audio
generate_audio

Whether to generate audio for the video.

selecttruetrue, false

Entrées

Fichiers que ce modèle accepte en plus du prompt.

EntréeAccepteFichiers max.
images
images
image4Obligatoire

Guide de prompt pour Vidu Q3

TARGET MODEL: Vidu Q3 (Shengshu — joint audio-video in one pass: lip-synced dialogue, voiceover, sound effects and music generated with the picture; 1–16 s; native camera control and "smart cuts" multi-shot). Plain descriptive prose. No negative prompt, no weighting syntax; the style and motion-strength parameters are documented as ineffective on Q3, so everything about look and energy is said in words.

Règles à respecter

  • Describe subject, action, setting, style, camera movement and mood as prose; the official example is one flowing paragraph ("In an ultra-realistic fashion photography style featuring light blue and pale amber tones, an astronaut in a spacesuit walks through the fog. The background consists of enchanting white and golden lights…"). No tag lists, no bracketed commands, no (word:1.3).
  • Look and motion energy are words, not parameters: name the visual style in the sentence ("anime", "ultra-realistic fashion photography") and the amount of motion ("barely moves", "sprints") — the style and movement_amplitude parameters do nothing on Q3.
  • Sound is generated with the picture whenever audio is on: write it as content — a spoken line in quotes attributed to a visible speaker, named sound effects tied to visible actions, the music's character. Multi-speaker conversation is supported: label each speaker uniquely and give each line its own sentence.
  • Dialogue languages that are supported: English, Japanese, Chinese. Another language may render poorly — keep such a line very short or carry the beat without on-camera speech.
  • One primary subject with supporting environmental detail per clip; do not name real people, brands or titles.
  • Hard cap 5,000 characters; a clip prompt is a paragraph, not a script.

Ce qui fonctionne le mieux

  • State the camera explicitly with film terms — pans, push-ins, tracking shots, "slow aerial drone shot", "static locked-off" — Q3 honors frame-level camera direction and a named move is more consistent than an implied one.
  • Multi-shot is native ("smart cuts"): when the story has beats, write them in order, each with its framing and its sound, and let the model place the cuts; a single-take clip says so ("one continuous shot").
  • Environmental sound follows environmental detail: "busy Tokyo crosswalk at night" pulls tires, train horns and signal beeps — describe the world and its sounds together ("campfire crackles… sparks drifting… crickets chirping, occasional owl hoot").
  • Match content to seconds: a 5-second clip is one beat; use the 16-second ceiling for a short sequence, not for a longer single action.

Issu du manuel LUVI pour cette famille : les mêmes règles qu'appliquent le réécriveur de l'espace de travail et le moteur de prompt MCP.

Autres modèles de cette famille