Which models generate sound with the video?
Seedance 2.0 and 2.5, Veo 3.1, Gemini Omni, MiniMax H3, HappyHorse, FLUX 3 Video, Grok Imagine Video and Kling 3 all produce the soundtrack in the same pass as the picture — dialogue, effects, ambience and music — but they differ in whether audio is on by default and whether it costs extra.
The joint audio-video families
| Family | Audio by default | Price effect | How to ask for sound |
|---|---|---|---|
| Seedance 2.0 / 2.5 | on | included | typed audio channels in the prompt |
| Veo 3.1 | off | standard ×2, Fast ×1.2, Lite has no audio | switch on generate_audio, describe the sound |
| Gemini Omni 1.1 | on, no switch | included | describe it in prose |
| MiniMax H3 | always on | included | dialogue tags and a soundscape line |
| HappyHorse 1.0 / 1.1 | on | included | must be named in words, else near-silent |
| FLUX 3 Video | on, switchable | included | a speaker must be visible for spoken lines |
| Grok Imagine Video | always on | included | name the sound, negation is ignored |
| Kling 3 (v3.0, O3) | on | varies by row, shown in the estimate | speaker labels for dialogue |
Every one of these families has a prompting guide on its model page; the audio syntax is the part that differs most between them.
Three things to know before you generate
- Silence is a parameter, not a sentence. Writing "no music" in the prompt does not mute a model; on Grok and H3 it is ignored, on others it can summon music. Use the switch where one exists.
- Dialogue needs a visible speaker on FLUX 3 Video and works best that way everywhere; a quoted line with nobody on screen can render as on-screen text.
- Veo ships silent unless you ask. Its audio switch is off on LUVI because it changes the price; the estimate updates when you turn it on.
Voice-first alternatives
If you need a specific voice, generate the speech with a text-to-speech model (ElevenLabs, Gemini TTS, MiniMax Speech, Grok Speech) and drive a lip-sync model with it. The audio's length then sets the price.
Frequently asked
Can I turn the sound off?
Only where the model has a switch: Veo and FLUX 3 Video have generate_audio. H3, Grok video and Gemini Omni always generate audio; on Seedance you leave the audio channel out of the prompt.
Why did my clip come back nearly silent?
Most of these models need the sound named in the prompt. HappyHorse in particular returns near-silence if no audio line is written; the family's prompting guide shows the syntax.
Can I add a real voice track instead?
Yes — lip-sync models such as OmniHuman, VEED and sync take your own audio file and animate the mouth to it.