MiniMax H3 vs Gemini Omni 1.1 vs Wan 3.0: three audio-video models, three prompt dialects
All three generate picture and sound in one pass, and all three read a prompt differently: H3 wants a fielded document, Omni wants a caption with a continuity clause, Wan wants an additive formula. Prices, limits and the one rule that breaks each of them.

Three video models arrived on LUVI within a week of each other in September, and they share the headline — sound generated together with the picture, dialogue included. They do not share a prompt language. We wrote a prompting guide for each from the vendors' own documentation; this is the comparison that fell out of that work.
The numbers
- Base price on LUVI. H3 from 2,080 credits for 8 s (H3 Max 1,216). Omni 1.1 Flash from 2,200 credits for 10 s. Wan 3.0 from 2,000 credits for 5 s (Prime 2,800).
- Native resolutions. H3: 480P, 768P, 2K. Omni: 360p, 720p; 1080p and 4K are Google's upscaler. Wan: 480p, 720p, 1080p.
- Clip length. H3: 4–15 s. Omni: 3–10 s, extendable to 40 s. Wan: up to 30 s.
- Modes on LUVI. H3: text, image, reference. Omni: text, image, reference, edit, extend. Wan: text, image, reference.
- Dialogue languages. H3: 11 (Turkish not among them). Omni: English evaluated, others may work. Wan: Chinese and English.
- Reference media. H3: up to 9 images, 3 clips, 3 audio, 12 files total. Omni: images and up to 3 short clips, no audio. Wan: up to 10 images, 5 clips, 5 audio, 20 files.
MiniMax H3: why is the prompt a document?
H3 does not want a caption. It wants three named fields — integrated_multimodal_description, overall_soundscape, non_diegetic_music — and it wants all three every time. Leave an audio field out and the model does not fall silent; it invents. Camera is prose built from twelve exact verbs and four modifiers ("The camera pushes in with small amplitude at slow speed toward…"), cuts are timestamped to the millisecond ("[Shot 2] At 00:03.500, the camera cuts to…"), dialogue sits in a tag with its language: <d>[English] I get off at the next station.</d>. A proper noun — a celebrity, a film, a brand — blocks the whole request. Bracketed camera commands from the older Hailuo line are dead syntax here.
Best at: scripted multi-shot with dialogue, and reference packs that lock a face and a voice at once.
Gemini Omni 1.1: why must a caption ask for a single take?
Omni is the opposite temperament: a 45-word caption in plain prose is the ideal prompt, and the model fills in the world. The trap is that Omni cuts by default — ask for a lakeside shot and you may get a three-shot montage of one. Every single-take prompt needs a continuity clause in the first sentence ("In a single unbroken scene…"). Sound goes in its own sentence behind "Sound design:", music has to be asked for, and dialogue uses a colon with no quotation marks — quotes burn the words into the frame as on-screen text. There is no negative field; exclusions are terminal fragments ("No dialogue. No text overlay on screen."). Edits are one clause plus "Keep everything else the same."
Best at: quick, cheap, natural-looking clips; conversational edits; extending a clip you already have.
Wan 3.0: what is the additive formula?
Alibaba publishes the prompt as a formula and the model reads it in that order: Entity + Scene + Motion, then Aesthetic control + Stylization, then Sound (voice = lines + emotion + tone + speed + timbre + accent; effect = source + action + ambience; music = score + style), then multi-shot as "Shot 1 [0–3 s] … Shot 2 [4–6 s] Hard cut transition, fixed camera…". Uploaded media is addressed by name and used as a noun: "The cat in Image 1 plays in the room in Image 2." Its own do-not list is short and worth memorising: no real people by name, no rapid scene changes inside one clip, no shot built on exact text or word-exact lip-sync.
Best at: long clips (30 s), mixed reference packs with audio, and Chinese-language dialogue.
The rule that breaks each one
- H3: omit an audio field. You get invented sound, not silence.
- Omni: forget the continuity clause. You get cuts.
- Wan: name a real person. The request fails.
Which one should I pick?
A dialogue scene with a locked cast: H3. A fast draft or an edit of something that exists: Omni. A long clip or a reference pack that includes a voice: Wan 3.0. All three guides are on the model pages — MiniMax H3, Gemini Omni, Wan — and the Workspace applies them when the guide toggle is on. For how the credit prices above are built, see how credits work.
Questions
Which one is cheapest?
For a short clip at the base resolution: Wan 3.0 from 2,000 credits (5 s), MiniMax H3 from 2,080 credits (8 s), Gemini Omni 1.1 from 2,200 credits (10 s). Per second, Omni is the cheapest of the three at its default.
Which one holds a character across shots?
All three take reference media. Wan 3.0 accepts up to 20 files including audio; H3 takes up to 12 with images, clips and voices; Omni takes images and short clips but no audio.
Which one can be edited after the fact?
Omni 1.1 has a video-edit row on LUVI, Wan 3.0 has one too (2.7 line); H3 does not — you re-generate.