PixVerse·by PixVerseReference to video

Pixverse v6 Reference-to-Video

Pixverse v6 Reference-to-Video model. High-quality video generation from image prompts.

Open in workspacepixverse/v6/reference-to-video

Parameters

ParameterTypeDefaultRange or options
duration
duration

The duration of the generated media in seconds (1-15).

select51, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15
aspect ratio
aspect_ratio

The aspect ratio of the generated video.

select16:916:9, 9:16, 1:1, 4:3, 3:4, 2:3, 3:2, 21:9
seed
seed

Random seed for reproducibility.

number
quality
quality

The resolution of the generated video.

select720p360p, 540p, 720p, 1080p
sound
sound

Whether sound is generated simultaneously when generating a video.

selecttruetrue, false

Inputs

Files this model takes in addition to the prompt.

InputAcceptsMax files
images
images
image4Required

Prompting guide for PixVerse

TARGET MODEL: PixVerse (AISphere video). Sequential parser — earlier tokens weigh most. Priority order is official: subject identity → action → environment → technical cues LAST.

Rules that must hold

  • Text-to-video is THREE sentences: S1 = subject + defining traits + ONE action + location. S2 = ONE camera move + named style/lens/lighting/composition cues. S3 = positive stability constraints. Sweet spot 50-80 words — past ~200 the prompt dilutes its own control.
  • ONE action, restrained — competing verbs produce artifacts. Show speed as physical evidence (motion blur, streaking lights); never repeat "fast".
  • ONE camera move per shot, named simply in prose ("slow macro push-in", "steady lateral tracking", "low-angle tilting upward").
  • KILL the praise words — PixVerse's own guide says "cinematic", "beautiful", "epic", "professional" sample too broadly. Replace each with a named physical cue: "warm rim light", "cyan-magenta contrast", "35mm", "anamorphic 2.39:1", "one-point perspective".
  • The prompt box takes POSITIVE constraints only: "Hands remain natural", "Cup shape remains stable", "Product silhouette stays intact". Naming a defect noun ("no bent fingers") can summon it — convert every exclusion to its positive form.
  • Don't write "realistic" as a style word — realism is PixVerse's unstyled base; named cues carry the look.
  • When the user wants sound, write the audio into the shot as content ("audio includes soft room tone and faint spoon clink") — never as a toggle or promise.

What works best

  • Reference shape (official, verbatim rhythm): "A ceramic coffee cup sits on a dark wooden table as steam rises in slow curls. Slow macro push-in, warm tungsten side light, shallow depth of field, quiet morning cafe background. Cup shape remains stable, no text overlay, audio includes soft room tone and faint spoon clink."
  • Character consistency is a vocabulary lock: keep the character's description a fixed paragraph (face, hair, outfit, build) and paste it WORD-FOR-WORD every shot — identical phrasing carries identity; drift in words is drift in face.
  • Technical cues ride the tail, never the lead.

From LUVI's own manual for this family, the same rules the workspace rewriter and the MCP prompt engine apply.

Other models in this family