HappyHorse·von AlibabaReference to video

HappyHorse-1.0 Reference-to-video

Generates videos from one to nine reference images and a text prompt, supporting 720P or 1080P output, flexible aspect ratios, and durations from 3 to 15 seconds.

Im Arbeitsbereich öffnenalibaba/happyhorse-1.0/reference-to-video

Parameter

ParameterTypStandardBereich oder Optionen
Duration (s)
duration

Video duration in seconds (3–15). Billed per second.

slider53 – 15
resolution
resolution

Output video resolution.

select1080P720P, 1080P
seed
seed

Random seed for video generation. Use -1 for a random seed.

number-1-1 – 2147483647
ratio
ratio

Aspect ratio of the generated video.

select16:916:9, 9:16, 1:1, 4:3, 3:4

Eingaben

Dateien, die dieses Modell zusätzlich zum Prompt entgegennimmt.

EingabeAkzeptiertMax. Dateien
images
images
image4Erforderlich

Prompt-Leitfaden für HappyHorse

TARGET MODEL: HappyHorse (Alibaba 快乐小马 — native joint audio-video, Kling-lineage craft). One pass renders the picture AND its sound together. The control surface is thin: no negative_prompt, no cfg/guidance, no weights — the prose is the whole instrument.

Regeln, die gelten müssen

  • Lead with the SUBJECT, not the action, and put the camera cue LAST where it weighs most: [subject] [does one action] in [setting], [time of day], [ONE atmosphere or camera cue].
  • 20-60 words for a single beat. Longer dilutes its own control — faces drift toward a generic average and hands lose geometry. Hard ceiling 2500 characters.
  • AUDIO MUST BE NAMED IN WORDS or the clip comes back near-silent — the model does not infer sound from what it sees. Every rewrite carries a sound line, written as content ("Audio: ..." / "Foley: ..."), never as a toggle or a promise. Three tiers, as the shot needs them: foreground (dialogue, hero SFX) · mid (action Foley — footsteps, sizzle, water, cloth) · background (ambience, room tone, music).
  • Dialogue is a quoted line plus a language tag: says calmly in French, "Tu as vu ca?" — and a non-English line is written in its NATIVE SCRIPT, never romanized (lip-sync is phoneme-level and trained on real script). Keep spoken lines short and the speaker front-facing. The prompt's own language does NOT drive the speech; the tag and the quoted script do.
  • Lip-sync covers seven languages — English, Mandarin, Cantonese, Japanese, Korean, German, French (1.1 widens the set). A line outside them syncs poorly: keep such a line very short, or carry the meaning without on-camera speech and put the language the user asked for in the audio bed instead of the mouth.
  • ONE cinematography cue — a lens OR a light recipe OR a camera move. Stacking five cancels them out. Keep the camera in its own clause, separate from the character's action, and state zoom, shake or cuts explicitly so camera motion is never read as subject motion.
  • ONE primary action, described cleanly. Over-detailing motion DEGRADES it: "A child runs through a field" outperforms the same line padded with hair, dust and arm-swing. Never micro-choreograph.
  • No weights, no cfg, no negative_prompt: booru tag-soup, JSON and weighted parentheses all underperform here. Exclusions ride inline, concrete and short, at the tail ("no music", "no dialogue", "no people in frame").
  • Kill the praise adjectives — beautiful, stunning, masterpiece, epic, hyperrealistic, ultra-detailed. Replace each with concrete capture language: "overcast daylight, wet asphalt, neon pink and cyan reflections in the puddles, 35mm telephoto, shallow depth of field".
  • One lighting recipe and one color story, not five. A director's name only when paired with the visual it stands for ("Wong Kar-wai — saturated greens and reds, telephoto compression"); the name alone underperforms.

Was am besten funktioniert

  • Multi-shot inside one clip is this model's differentiator, and unlike most video models its TIMECODES ARE HONORED. When the user's intent has beats, cuts or a small story, write it as: Shot N (framing, start-end s): [camera] [subject action] [light] [audio] — about 5 seconds per shot, identity holding across the cuts.
  • Official rhythm to imitate for a single beat: "A young woman in a red coat walks down a wet city street at night, neon reflections."
  • Multi-speaker is character index plus timecode: 0-4s: character1 says in French, "..."; 4-8s: character2 laughs and replies, "...".
  • Keep shots short and physics scripted. Long takes lose spatial consistency; unscripted physics breaks (doors that close themselves, wrong water speed); fast action smears fine wardrobe detail; slow motion is modest here, not dramatic time dilation.
  • Markdown sections parse cleanly for a single continuous take: ## Subject / ## Action / ## Camera / ## Lighting / ## Audio.

Aus LUVIs eigenem Handbuch für diese Familie; dieselben Regeln wendet der Rewriter im Arbeitsbereich und die MCP-Prompt-Engine an.

Weitere Modelle dieser Familie