Vidu Q3·de ViduText to video

Vidu Q3-Pro Text-to-video

Vidu Q3-Pro Text-to-Video is an advanced AI video generation model that creates high-quality videos directly from text descriptions. With support for multiple styles, resolutions up to 1080p, and optional audio generation, it delivers cinematic results with smooth motion and rich detail.

Abrir en el espacio de trabajovidu/q3-pro/text-to-video

Parámetros

ParámetroTipoPor defectoRango u opciones
duration
duration

The duration of the generated media in seconds.

slider51 – 16
aspect ratio
aspect_ratio

The aspect ratio of the generated media.

select4:316:9, 9:16, 4:3, 3:4, 1:1
resolution
resolution

The resolution of the generated media.

select720p540p, 720p, 1080p
seed
seed

The random seed to use for the generation. -1 means a random seed will be used.

number
bgm
bgm

The background music for generating the output.

selecttruetrue, false
generate audio
generate_audio

Whether to generate audio.

selecttruetrue, false

Entradas

Sin archivos de referencia: este modelo trabaja solo con el prompt.

Guía de prompts para Vidu Q3

TARGET MODEL: Vidu Q3 (Shengshu — joint audio-video in one pass: lip-synced dialogue, voiceover, sound effects and music generated with the picture; 1–16 s; native camera control and "smart cuts" multi-shot). Plain descriptive prose. No negative prompt, no weighting syntax; the style and motion-strength parameters are documented as ineffective on Q3, so everything about look and energy is said in words.

Reglas que deben cumplirse

  • Describe subject, action, setting, style, camera movement and mood as prose; the official example is one flowing paragraph ("In an ultra-realistic fashion photography style featuring light blue and pale amber tones, an astronaut in a spacesuit walks through the fog. The background consists of enchanting white and golden lights…"). No tag lists, no bracketed commands, no (word:1.3).
  • Look and motion energy are words, not parameters: name the visual style in the sentence ("anime", "ultra-realistic fashion photography") and the amount of motion ("barely moves", "sprints") — the style and movement_amplitude parameters do nothing on Q3.
  • Sound is generated with the picture whenever audio is on: write it as content — a spoken line in quotes attributed to a visible speaker, named sound effects tied to visible actions, the music's character. Multi-speaker conversation is supported: label each speaker uniquely and give each line its own sentence.
  • Dialogue languages that are supported: English, Japanese, Chinese. Another language may render poorly — keep such a line very short or carry the beat without on-camera speech.
  • One primary subject with supporting environmental detail per clip; do not name real people, brands or titles.
  • Hard cap 5,000 characters; a clip prompt is a paragraph, not a script.

Lo que mejor funciona

  • State the camera explicitly with film terms — pans, push-ins, tracking shots, "slow aerial drone shot", "static locked-off" — Q3 honors frame-level camera direction and a named move is more consistent than an implied one.
  • Multi-shot is native ("smart cuts"): when the story has beats, write them in order, each with its framing and its sound, and let the model place the cuts; a single-take clip says so ("one continuous shot").
  • Environmental sound follows environmental detail: "busy Tokyo crosswalk at night" pulls tires, train horns and signal beeps — describe the world and its sounds together ("campfire crackles… sparks drifting… crickets chirping, occasional owl hoot").
  • Match content to seconds: a 5-second clip is one beat; use the 16-second ceiling for a short sequence, not for a longer single action.

Del manual propio de LUVI para esta familia: las mismas reglas que aplican el reescritor del espacio de trabajo y el motor de prompts de MCP.

Otros modelos de esta familia