Which models generate sound with the video?

Seedance 2.0 and 2.5, Veo 3.1, Gemini Omni, MiniMax H3, HappyHorse, FLUX 3 Video, Grok Imagine Video and Kling 3 all produce the soundtrack in the same pass as the picture — dialogue, effects, ambience and music — but they differ in whether audio is on by default and whether it costs extra.

Last updated:

The joint audio-video families

FamilyAudio by defaultPrice effectHow to ask for sound
Seedance 2.0 / 2.5onincludedtyped audio channels in the prompt
Veo 3.1offstandard ×2, Fast ×1.2, Lite has no audioswitch on generate_audio, describe the sound
Gemini Omni 1.1on, no switchincludeddescribe it in prose
MiniMax H3always onincludeddialogue tags and a soundscape line
HappyHorse 1.0 / 1.1onincludedmust be named in words, else near-silent
FLUX 3 Videoon, switchableincludeda speaker must be visible for spoken lines
Grok Imagine Videoalways onincludedname the sound, negation is ignored
Kling 3 (v3.0, O3)onvaries by row, shown in the estimatespeaker labels for dialogue

Every one of these families has a prompting guide on its model page; the audio syntax is the part that differs most between them.

Three things to know before you generate

  1. Silence is a parameter, not a sentence. Writing "no music" in the prompt does not mute a model; on Grok and H3 it is ignored, on others it can summon music. Use the switch where one exists.
  2. Dialogue needs a visible speaker on FLUX 3 Video and works best that way everywhere; a quoted line with nobody on screen can render as on-screen text.
  3. Veo ships silent unless you ask. Its audio switch is off on LUVI because it changes the price; the estimate updates when you turn it on.

Voice-first alternatives

If you need a specific voice, generate the speech with a text-to-speech model (ElevenLabs, Gemini TTS, MiniMax Speech, Grok Speech) and drive a lip-sync model with it. The audio's length then sets the price.

Frequently asked

Can I turn the sound off?

Only where the model has a switch: Veo and FLUX 3 Video have generate_audio. H3, Grok video and Gemini Omni always generate audio; on Seedance you leave the audio channel out of the prompt.

Why did my clip come back nearly silent?

Most of these models need the sound named in the prompt. HappyHorse in particular returns near-silence if no audio line is written; the family's prompting guide shows the syntax.

Can I add a real voice track instead?

Yes — lip-sync models such as OmniHuman, VEED and sync take your own audio file and animate the mouth to it.

More questions