MiniMax H3·par MiniMaxImage to videoNouveau
MiniMax H3 Image-to-Video
Animate a still image into a 24fps clip with synchronized audio with MiniMax H3 — native 2K. Optional end frame for start-to-end motion control. 480P, 768P, 2K; 4-15 seconds. Aspect ratio follows the first frame.
minimax/h3/image-to-videoParamètres
| Paramètre | Type | Par défaut | Plage ou options |
|---|---|---|---|
Duration (s) durationVideo duration in seconds (4-15). | select | 8 | 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 |
Resolution resolutionVideo resolution. Per-second price: 480P $0.038/s · 768P $0.08/s · 2K $0.13/s. | select | 2K | 480P (480P), 768P (768P), 2K (2K) |
Aspect Ratio ratioAspect ratio follows the first frame. | select | adaptive | adaptive |
Entrées
Fichiers que ce modèle accepte en plus du prompt.
| Entrée | Accepte | Fichiers max. | |
|---|---|---|---|
First Frame image | image | 1 | Obligatoire |
Last Frame end_image | image | 1 | Facultatif |
Guide de prompt pour MiniMax H3
TARGET MODEL: MiniMax H3 (V2 endpoint — H3 and H3-Max). Joint audio-video in one pass: voice, sound effects and music are modeled together and are ALWAYS generated; there is no mute flag, no negative_prompt, no weighting syntax, no seed. The prompt is a fielded document in H3's own grammar, not a caption — the rewriter's job is to emit that document.
Règles à respecter
- Emit the fielded structure, lowercase field names, colon-terminated, one blank line between fields: integrated_multimodal_description: … / overall_soundscape: … / non_diegetic_music: … — in that order, every time.
- ALWAYS emit BOTH audio fields. An omitted field is filled with invented sound, never silence. overall_soundscape = 1-4 sentences in one paragraph: ambience, physical action sounds, non-verbal human sounds — no dialogue, no diegetic music. non_diegetic_music = 1-3 sentences on instrumentation, tempo, rhythm and dynamic change — never mood words, never the emotional function of the score. Silence is the literal N/A (music N/A is reliable; soundscape N/A only when the user asks for a fully silent clip).
- Camera is PROSE, never a bracket. Twelve moves, exact spelling: Zoom In/Out · Push In/Pull Out · Pan Left/Right · Truck Left/Right · Tilt Up/Down · Pedestal Up/Down · Arc Shot · Tracking Shot · Static Shot · Shake Slightly/Strongly · POV · Roll Clockwise/Counterclockwise. Four modifier literals only: with small amplitude · with large amplitude · at slow speed · at fast speed (medium amplitude / normal speed are expressed by omission). Form: "The camera pushes in with small amplitude at slow speed toward the folded letter in her hands." Never [Push in], [pan], [static], "dolly in", ECU/MCU/OTS, or lens/f-stop numbers — write "shallow depth of field" and lowercase shot sizes ("a medium-wide shot frames…").
- One shot, one move, one visible action, one end frame. A second camera move costs a cut, not a comma: "[Shot 2] At 00:03.500, the camera cuts to …" — no timestamp on Shot 1, strictly increasing MM:SS.mmm cut times inside the clip, cut verbs from the allowlist only (the camera cuts to · the shot cuts to · the shot transitions to · the shot changes to · the shot switches to). Cross-dissolves and fades only when the user asked; on a hard-cut clip carry a standing "No dissolves, no morph transitions" clause because the model drifts toward soft transitions.
- Dialogue is a tagged block: The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d> — delivery, action and speaker ID sit OUTSIDE the tag; only the language tag and the verbatim, untranslated words sit inside. Speaker IDs (S1, S2) stay stable across shots. Speech truncated by the clip end gets <cutoff>. Voiceover requires the exact phrase "says in an off-screen voiceover" plus an immediate "lips remain closed" clause or the character lip-flaps.
- Stable dialogue languages: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish. Turkish is OUTSIDE the set: never invent a language tag — write the shot with no <d> block, describe the performance physically ("she speaks two short sentences to camera, lips moving with measured phrasing, no audible dialogue directed to the microphone"), set the room in overall_soundscape, and let the user lay ADR in post.
- Proper nouns BLOCK the request (character names, film titles, studios, celebrities trip input moderation). Describe the silhouette; a director's or auteur's STYLE name is safe ("the Hitchcock camera movement").
- "Cinematic" is legal exactly once, in the style slot at the head of [Shot 1] ("[Shot 1] Cinematic, live-action, a medium-wide shot frames …"); elsewhere it and "beautiful" are banned abstractions. Styles that work: Cinematic, live-action, 2D-animated, 3D CG, claymation, watercolor, vintage film.
- On-screen text goes in ASCII double quotes, verbatim, untranslated: A red neon sign reading "OPEN" glows above the doorway. Unquoted text returns as letter-shaped noise; one single-line caption on screen at a time.
- No symbols in the prose — no arrows, slashes, plus signs or connectors; the model paints them into the frame. No weights, no (word:1.3), no BREAK, no tag-soup.
- Exclusions are positive end-states first, and a short trailing "No X, no Y" clause only as backstop (MiniMax's own templates end "no split screen, no hard cuts, no random camera shake, no watermark").
- Keep faces medium or closer; a wide frontal face degrades regardless of resolution (vendor-acknowledged). Wide frames get back or rear-three-quarter views. Max three characters with on-screen action or dialogue per shot.
Ce qui fonctionne le mieux
- Target 350-500 words of body; hard ceiling 7,000 characters.
- Two shots by default, three when the beats demand it; spend a cut only for new information (subject, space, state, viewpoint, time) — a change of distance is a camera move, not a cut. Beat budget: 5 s → 3-4 beats · 10 s → 5-7 · 15 s → 6-9, one primary action per beat.
- Dialogue density: a sentence or two per five seconds. Over-write and the delivery accelerates; under-write and the model fills the gap with invented speech. Two lines of dialogue in 15 s is the practical ceiling before background music silently vanishes.
- Diegetic sound (radio, busker, phone speaker, on-camera band) is written in the body next to the picture that motivates it; score goes in non_diegetic_music — a routing rule, not a style choice.
- Fast action shatters limbs and clothing: one action per shot, "with small amplitude at slow speed", and cut rather than accelerate.
- Timestamps address picture reliably and sound unreliably: schedule cuts, never a music hit against a frame.
Issu du manuel LUVI pour cette famille : les mêmes règles qu'appliquent le réécriveur de l'espace de travail et le moteur de prompt MCP.