Wan·AlibabaImage to videoYeni

Wan 3.0 Image-to-Video

Bir ilk kareyi Alibaba Wan 3.0 ile senkron sesli videoya dönüştürün; isteğe bağlı son kare ile hareketin nerede biteceğini belirleyin. 2-30 saniye; 480p, 720p, 1080p. En-boy oranı ilk kareyi takip eder.

En uygun olduğu işler

Fotoğraf animasyonuİlk-son kare kontrolüÜrün çekimleriSesli video üretimi
Çalışma alanında açalibaba/wan-3.0/image-to-video

Parametreler

ParametreTürVarsayılanAralık veya seçenekler
Süre (sn)
duration

Video süresi (saniye, 2-30). Çıktının saniyesi başına ücretlendirilir.

select52, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30
Çözünürlük
resolution

Çıktı çözünürlüğü. Saniye başı fiyatlanır: 720p yaklaşık 2×, 1080p yaklaşık 4× 480p fiyatı.

select1080p480p (480p), 720p (720p), 1080p (1080p)
Ses Üret
audio

Çıktı videosunun senkron ses içerip içermediği. Fiyat iki durumda da aynı.

selecttruetrue, false

Girdiler

Prompt'a ek olarak bu modelin aldığı dosyalar.

GirdiKabul ederEn fazla dosya
İlk Kare
image
image1Zorunlu
Son Kare
last_image
image1İsteğe bağlı

Wan için prompt rehberi

TARGET MODEL: Wan (Alibaba 万相 — image and joint audio-video generation, 2.5 → 3.0). Alibaba publishes the prompt as an additive formula and the model reads it in that order: Entity + Scene + Motion (basic) → + Aesthetic control + Stylization (advanced) → + Sound description (voice / sound effect / background music) → multi-shot as Overall description + Shot number + Timestamp + Shot content. No weighting syntax, no tag-soup; the video rows have no negative prompt at all (write exclusions as "No dialogue." / "No background music." style sentences), and expansion is a parameter, not a word.

Uyulması gereken kurallar

  • Keep the formula order: Entity (appearance — adjectives or short phrases) → Scene (environment) → Motion (amplitude, speed, effect) → Aesthetic control (light source, lighting, shot size, camera angle, lens, camera movement) → Stylization (one style word: cyberpunk, line-art illustration, claymation, felt, pixel, 3D cartoon, black-and-white animation) → Sound. Describe, never enumerate keywords.
  • Audio is generated natively; write it as content. Voice = the character's lines + emotion + tone + speed + timbre + accent ("says in a relaxed tone, at a moderate speed, with a clear bright voice, in American English: "…""). Sound effect = source material + action + ambient sound ("a small glass ball falls from the table onto the wooden floor with a 'thud' in a quiet room"). Background music = score + style ("suspenseful background music"). Lines go in double quotes; on-screen text also goes in double quotes.
  • Multi-character dialogue: give every speaker a UNIQUE label (never a pronoun or a synonym), anchor each line to a visible action, give each voice its own tone/emotion, and order the lines in time ("first …, then …").
  • Multi-shot is written as: one overall description line (story, perspective, the subject's look — repeat the subject here, not inside every shot), then "Shot 1 [0–3 s] …", "Shot 2 [4–6 s] Hard cut transition, fixed camera position, …". Timestamps are honored; 4–6 seconds per shot; state the transition ("hard cut") and the framing of every shot. The labeled storyboard form is also official: "(0:00 - 0:03) Camera: medium shot. Scene: … Action: … Atmosphere: …".
  • Name the framing and the camera treatment of every shot ("static shot" / "fixed shot" to hold still; slow push-in, pull-out, tracking shot, orbit, pan, tilt, drone fly-through). Left unnamed, Wan edits on your behalf and cuts inside a single clip. One main camera action per shot.
  • Never: a real person by name · rapid scene changes inside one clip · exact text legibility as the point of the shot · very long or complex action chains · a shot that depends on word-exact lip-sync. Quoted lines are performed and synced, but do not build the beat on a specific syllable landing on a frame.
  • Write in English (Chinese and English are the supported languages); a non-English line stays inside its quotes, untranslated, with the accent named in the voice description.
  • Keep the whole prompt under 2,000 characters so it ports across every Wan row (the 2.6 image rows truncate at 2,000; 2.7 rows at 5,000).

En iyi sonuç verenler

  • The cinematic vocabulary the guide itself uses: light source (daylight, firelight, overcast light, clear-sky light) · light type (soft, hard, side light, high contrast) · time (daytime, night, dawn) · shot size (close-up, close shot, wide angle, extreme full shot) · angle (over-the-shoulder, high angle, aerial) · shot type (clean single, two-shot, group shot, establishing shot) · composition (center, left-heavy) · lens (long-focus, ultra-wide fisheye) · tone (warm tones, low saturation) · effects (tilt-shift, time-lapse).
  • Camera moves carry intent: push-in for intimacy or tension, pull-out to reveal scale or isolation, tracking to walk beside the subject, orbit to crown the subject, fixed camera for stillness and focus.
  • Match beats to seconds: roughly 5–8 seconds per beat; a 5-second clip carries about 2.5 seconds of actual speech. Under-write and the model fills the gap; over-write and the delivery rushes.
  • Verbs beat adjectives for sound ("the soft haptic tick of a thumb tapping the screen" over "realistic phone sounds"); place layers in depth ("foreground clink", "distant chatter", "off-camera door closing"); name the tempo when motion must sit on a beat ("driving bass at 100 BPM, each strike landing on the beat").
  • Suppress by sentence, not by field: "No dialogue." "No background music." "No on-screen text, no subtitles." at the tail.

LUVI'nin bu aile için yazdığı manuelden; çalışma alanındaki yeniden yazıcı ve MCP prompt motoru aynı kuralları uygular.

Bu ailedeki diğer modeller