How to prompt HappyHorse: name the sound, or the clip comes back silent

HappyHorse generates picture and sound together, but it does not infer sound from what it sees: a prompt with no audio line returns a near-silent video. It is also one of the few video models that honors timecoded shots, which is the opposite of the Seedance 2.5 rule.

Published:

Four widescreen frames sit along a glowing timeline in a dark studio, amber tick marks between them, and warm rings of sound rise from the second frame while the others carry flat dim lines.

HappyHorse is Alibaba's joint audio-video model—picture, dialogue, effects and ambience come out of one pass, with phoneme-level lip-sync. On LUVI it runs as text-to-video, image-to-video, reference-to-video and video edit, in 720P or 1080P, from three to fifteen seconds. It is billed per second, so the credit estimate appears in the Workspace before you run, and 1080P costs more per second than 720P. Two of its rules appear nowhere else in our manuals, and both of them are the reason people give up on the model too early.

Why did my HappyHorse clip come back almost silent?

HappyHorse does not infer sound from what it sees. A prompt that describes a thunderstorm perfectly well, without a sentence naming the sound, returns a near-silent video—because the audio is co-generated from words, not from the picture. This catches people because the vendor's own text-to-video API reference contains no audio guidance at all: its example prompt describes only visual elements, a cardboard city with a train passing through.

Our manual for the family therefore requires a sound line on every rewrite, written as content rather than as a promise. Use the three tiers the shot actually needs:

  • Foreground: dialogue and the one hero effect the scene is about.
  • Mid: action Foley tied to something visible—footsteps, sizzle, water, cloth.
  • Background: ambience, room tone, music.

Write it as its own clause: Audio: rain on the awning, distant traffic, no music. Exclusions are allowed inline at the tail and they work here—"no music", "no dialogue", "no people in frame".

How do you write dialogue it can lip-sync?

Give the line a language tag, quote it, and write a non-English line in its own script. The lip-sync is phoneme-level and trained on real script, so romanizing a line ("Der letzte Bus ist schon weg" written as sounds) syncs worse than writing it properly. The prompt's own language does not drive the speech—the tag and the quoted words do.

A young woman in a red coat stands under a shop awning at night and looks straight into the camera, neon reflections on the wet pavement, one slow push-in. She says in German, "Der letzte Bus ist schon weg." Audio: rain on the awning, distant traffic, no music.

Which languages actually lip-sync is where the public sources disagree. LUVI's manual for the family lists seven for 1.0—English, Mandarin, Cantonese, Japanese, Korean, German and French—and records that 1.1 widens the set. The fal model page for Happy Horse 1.1 names "English, French, Spanish, Turkish, Japanese, and more". Our checker follows the manual, so treat anything outside those seven as unproven: keep such a line very short, or carry the meaning without on-camera speech and put the language in the audio bed instead of in the mouth.

For more than one speaker, index the characters and give each turn its own time range: 0-4s: character1 says in French, "..."; 4-8s: character2 laughs and replies, "...". Keep the speaker front-facing and the lines short.

Do timecodes work on HappyHorse?

Yes, and this is the rule that makes HappyHorse worth learning. Timecoded shot beats are honored here, which is the opposite of what Seedance 2.5 does with the same syntax—that family paces by shot order and its per-second marks are unstable, as we covered in how to prompt Seedance 2.5. Carrying one habit into the other model is the most common way to waste a render on either.

Write beats as Shot N (framing, start-end s): and keep each one around five seconds, with identity holding across the cuts.

Shot 1 (wide establishing, 0-5s): A fisherman in a yellow oilskin coat hauls a net over the rail of a small boat, overcast dawn light. Audio: gulls, wind, water against the hull. Shot 2 (medium, 5-10s): The same fisherman drinks from a steel flask in the wheelhouse, locked-off camera. Audio: engine hum under the cabin, the flask knocking against the console.

Keep the shots short and the physics scripted. Long takes lose spatial consistency, unscripted physics breaks in familiar ways—doors that close themselves, water moving at the wrong speed—and fast action smears fine wardrobe detail.

How long can a HappyHorse prompt be?

The vendor's limit is generous and it does not fail loudly. Alibaba Cloud's text-to-video API reference states: "Maximum 5,000 non-Chinese characters or 2,500 Chinese characters. Excess is truncated." So an over-long English prompt is not rejected—the tail is silently cut, which is worse, because the closing lines are usually where people put the audio.

LUVI's checker warns earlier, at 2,500 characters, which is the conservative reading of that same limit. In practice the ceiling that matters is much lower: a single beat wants 20 to 60 words, because longer prompts dilute their own control—faces drift toward a generic average and hands lose geometry.

What makes a HappyHorse prompt worse?

Four habits, all of them imported from image models, and all of them measurable in the checker.

  • Praise adjectives. "Masterpiece", "ultra-detailed" and "hyperrealistic" are dead tokens: inert at best, degrading at worst. Replace them with capture language—"overcast daylight, wet asphalt, neon pink and cyan reflections in the puddles, 35mm telephoto, shallow depth of field".
  • Stacked camera cues. One cinematography cue per shot: a lens, or a light recipe, or a camera move. Five at once cancel each other out. Keep the camera in its own clause so its motion is never read as the subject's.
  • Over-choreographed motion. One primary action, described cleanly. "A child runs through a field" outperforms the same line padded with hair, dust and arm-swing.
  • Weights and ratios. There is no cfg and no weighting syntax, and the aspect ratio is a parameter with nine values—writing "16:9" into the prompt does nothing.

How we checked

  • Models: alibaba/happyhorse-1.1/text-to-video, /image-to-video and /reference-to-video, plus alibaba/happyhorse-1.0/video-edit.
  • Settings: three to fifteen seconds, 720P and 1080P, the nine aspect ratios the schema exposes, reference-to-video accepting up to nine images.
  • Date and what was read: September 22, 2026. Model schemas and credit estimates from LUVI, the prompting guide LUVI runs for this family, Alibaba Cloud's own API reference, and the fal model page linked above.
  • Results: the schema and the vendor reference agree on duration (3–15 s, default 5) and on all nine aspect ratios. The prompt limit quoted above is the vendor's wording. The language lists from the two public sources do not match, and the post says so rather than picking one.
  • Drawbacks: No outputs were generated for this post, so it makes no claim about output quality. Credit estimates move when a provider reprices a model—read the estimate in the Workspace before a long clip.

In LUVI

Pick HappyHorse in the Workspace and press the feather button before you generate. That is LuviTransLex, the free check that applies this family's prompting guide.

It is blunt with this model, because HappyHorse has a published list of words to avoid. Run A beautiful woman walking in a stunning city, masterpiece, ultra-detailed, hyperrealistic, epic cinematic lighting, 16:9, (neon:1.3), slow motion handheld orbit push-in with a zoom and a shake through it and you get two rewrites and four warnings: "masterpiece", "ultra-detailed" and "hyperrealistic" are deleted outright, "beautiful", "stunning" and "epic" are flagged as dead quality tags, the ratio is flagged as a parameter, and the weight is dropped while neon is kept. What survives is still a stacked camera instruction, which the guide tells you to cut to one cue by hand.

Reference-to-video takes up to nine images, so it is the mode to use when a character has to stay recognizable; address the supplied subjects as character1 and character2 and bind identity in words. You can also drive all of it from Claude or ChatGPT through the LUVI connector.

New here? Create a LUVI account and try a three-second clip first—it is the cheapest way to hear whether your audio line landed.

Sources

More from Guides