MiniMax H3·by MiniMaxImage to videoNew

MiniMax H3 Image-to-Video

Animate a still image into a 24fps clip with synchronized audio with MiniMax H3 — native 2K. Optional end frame for start-to-end motion control. 480P, 768P, 2K; 4-15 seconds. Aspect ratio follows the first frame.

Open in workspaceminimax/h3/image-to-video

Parameters

ParameterTypeDefaultRange or options
Duration (s)
duration

Video duration in seconds (4-15).

select84, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15
Resolution
resolution

Video resolution. Per-second price: 480P $0.038/s · 768P $0.08/s · 2K $0.13/s.

select2K480P (480P), 768P (768P), 2K (2K)
Aspect Ratio
ratio

Aspect ratio follows the first frame.

selectadaptiveadaptive

Inputs

Files this model takes in addition to the prompt.

InputAcceptsMax files
First Frame
image
image1Required
Last Frame
end_image
image1Optional

Prompting guide for MiniMax H3

TARGET MODEL: MiniMax H3 (V2 endpoint — H3 and H3-Max). Joint audio-video in one pass: voice, sound effects and music are modeled together and are ALWAYS generated; there is no mute flag, no negative_prompt, no weighting syntax, no seed. The prompt is a fielded document in H3's own grammar, not a caption — the rewriter's job is to emit that document.

Rules that must hold

  • Emit the fielded structure, lowercase field names, colon-terminated, one blank line between fields: integrated_multimodal_description: … / overall_soundscape: … / non_diegetic_music: … — in that order, every time.
  • ALWAYS emit BOTH audio fields. An omitted field is filled with invented sound, never silence. overall_soundscape = 1-4 sentences in one paragraph: ambience, physical action sounds, non-verbal human sounds — no dialogue, no diegetic music. non_diegetic_music = 1-3 sentences on instrumentation, tempo, rhythm and dynamic change — never mood words, never the emotional function of the score. Silence is the literal N/A (music N/A is reliable; soundscape N/A only when the user asks for a fully silent clip).
  • Camera is PROSE, never a bracket. Twelve moves, exact spelling: Zoom In/Out · Push In/Pull Out · Pan Left/Right · Truck Left/Right · Tilt Up/Down · Pedestal Up/Down · Arc Shot · Tracking Shot · Static Shot · Shake Slightly/Strongly · POV · Roll Clockwise/Counterclockwise. Four modifier literals only: with small amplitude · with large amplitude · at slow speed · at fast speed (medium amplitude / normal speed are expressed by omission). Form: "The camera pushes in with small amplitude at slow speed toward the folded letter in her hands." Never [Push in], [pan], [static], "dolly in", ECU/MCU/OTS, or lens/f-stop numbers — write "shallow depth of field" and lowercase shot sizes ("a medium-wide shot frames…").
  • One shot, one move, one visible action, one end frame. A second camera move costs a cut, not a comma: "[Shot 2] At 00:03.500, the camera cuts to …" — no timestamp on Shot 1, strictly increasing MM:SS.mmm cut times inside the clip, cut verbs from the allowlist only (the camera cuts to · the shot cuts to · the shot transitions to · the shot changes to · the shot switches to). Cross-dissolves and fades only when the user asked; on a hard-cut clip carry a standing "No dissolves, no morph transitions" clause because the model drifts toward soft transitions.
  • Dialogue is a tagged block: The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d> — delivery, action and speaker ID sit OUTSIDE the tag; only the language tag and the verbatim, untranslated words sit inside. Speaker IDs (S1, S2) stay stable across shots. Speech truncated by the clip end gets <cutoff>. Voiceover requires the exact phrase "says in an off-screen voiceover" plus an immediate "lips remain closed" clause or the character lip-flaps.
  • Stable dialogue languages: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish. Turkish is OUTSIDE the set: never invent a language tag — write the shot with no <d> block, describe the performance physically ("she speaks two short sentences to camera, lips moving with measured phrasing, no audible dialogue directed to the microphone"), set the room in overall_soundscape, and let the user lay ADR in post.
  • Proper nouns BLOCK the request (character names, film titles, studios, celebrities trip input moderation). Describe the silhouette; a director's or auteur's STYLE name is safe ("the Hitchcock camera movement").
  • "Cinematic" is legal exactly once, in the style slot at the head of [Shot 1] ("[Shot 1] Cinematic, live-action, a medium-wide shot frames …"); elsewhere it and "beautiful" are banned abstractions. Styles that work: Cinematic, live-action, 2D-animated, 3D CG, claymation, watercolor, vintage film.
  • On-screen text goes in ASCII double quotes, verbatim, untranslated: A red neon sign reading "OPEN" glows above the doorway. Unquoted text returns as letter-shaped noise; one single-line caption on screen at a time.
  • No symbols in the prose — no arrows, slashes, plus signs or connectors; the model paints them into the frame. No weights, no (word:1.3), no BREAK, no tag-soup.
  • Exclusions are positive end-states first, and a short trailing "No X, no Y" clause only as backstop (MiniMax's own templates end "no split screen, no hard cuts, no random camera shake, no watermark").
  • Keep faces medium or closer; a wide frontal face degrades regardless of resolution (vendor-acknowledged). Wide frames get back or rear-three-quarter views. Max three characters with on-screen action or dialogue per shot.

What works best

  • Target 350-500 words of body; hard ceiling 7,000 characters.
  • Two shots by default, three when the beats demand it; spend a cut only for new information (subject, space, state, viewpoint, time) — a change of distance is a camera move, not a cut. Beat budget: 5 s → 3-4 beats · 10 s → 5-7 · 15 s → 6-9, one primary action per beat.
  • Dialogue density: a sentence or two per five seconds. Over-write and the delivery accelerates; under-write and the model fills the gap with invented speech. Two lines of dialogue in 15 s is the practical ceiling before background music silently vanishes.
  • Diegetic sound (radio, busker, phone speaker, on-camera band) is written in the body next to the picture that motivates it; score goes in non_diegetic_music — a routing rule, not a style choice.
  • Fast action shatters limbs and clothing: one action per shot, "with small amplitude at slow speed", and cut rather than accelerate.
  • Timestamps address picture reliably and sound unreliably: schedule cuts, never a music hit against a frame.

From LUVI's own manual for this family, the same rules the workspace rewriter and the MCP prompt engine apply.

Other models in this family