MiniMax H3·by MiniMaxImage to videoNew
MiniMax H3 Fast Image-to-Video
Animate a still image into a 24fps clip with synchronized audio with MiniMax H3 Fast — the quickest, most economical 480P tier for rapid iteration. Optional end frame for start-to-end motion control. 480P; 5-15 seconds. Aspect ratio follows the first frame.
minimax/h3-fast/image-to-videoParameters
| Parameter | Type | Default | Range or options |
|---|---|---|---|
Duration (s) durationVideo duration in seconds (5-15). | select | 8 | 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 |
Resolution resolutionVideo resolution. Per-second price: 480P $0.0437/s. | select | 480P | 480P (480P) |
Aspect Ratio ratioAspect ratio follows the first frame. | select | adaptive | adaptive |
Inputs
Files this model takes in addition to the prompt.
| Input | Accepts | Max files | |
|---|---|---|---|
First Frame image | image | 1 | Required |
Last Frame end_image | image | 1 | Optional |
Prompting guide for MiniMax H3
TARGET MODEL: MiniMax H3 (V2 endpoint — H3 and H3-Max). Joint audio-video in one pass: voice, sound effects and music are modeled together and are ALWAYS generated; there is no mute flag, no negative_prompt, no weighting syntax, no seed. The prompt is a fielded document in H3's own grammar, not a caption — the rewriter's job is to emit that document.
Rules that must hold
- Emit the fielded structure, lowercase field names, colon-terminated, one blank line between fields: integrated_multimodal_description: … / overall_soundscape: … / non_diegetic_music: … — in that order, every time.
- ALWAYS emit BOTH audio fields. An omitted field is filled with invented sound, never silence. overall_soundscape = 1-4 sentences in one paragraph: ambience, physical action sounds, non-verbal human sounds — no dialogue, no diegetic music. non_diegetic_music = 1-3 sentences on instrumentation, tempo, rhythm and dynamic change — never mood words, never the emotional function of the score. Silence is the literal N/A (music N/A is reliable; soundscape N/A only when the user asks for a fully silent clip).
- Camera is PROSE, never a bracket. Twelve moves, exact spelling: Zoom In/Out · Push In/Pull Out · Pan Left/Right · Truck Left/Right · Tilt Up/Down · Pedestal Up/Down · Arc Shot · Tracking Shot · Static Shot · Shake Slightly/Strongly · POV · Roll Clockwise/Counterclockwise. Four modifier literals only: with small amplitude · with large amplitude · at slow speed · at fast speed (medium amplitude / normal speed are expressed by omission). Form: "The camera pushes in with small amplitude at slow speed toward the folded letter in her hands." Never [Push in], [pan], [static], "dolly in", ECU/MCU/OTS, or lens/f-stop numbers — write "shallow depth of field" and lowercase shot sizes ("a medium-wide shot frames…").
- One shot, one move, one visible action, one end frame. A second camera move costs a cut, not a comma: "[Shot 2] At 00:03.500, the camera cuts to …" — no timestamp on Shot 1, strictly increasing MM:SS.mmm cut times inside the clip, cut verbs from the allowlist only (the camera cuts to · the shot cuts to · the shot transitions to · the shot changes to · the shot switches to). Cross-dissolves and fades only when the user asked; on a hard-cut clip carry a standing "No dissolves, no morph transitions" clause because the model drifts toward soft transitions.
- Dialogue is a tagged block: The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d> — delivery, action and speaker ID sit OUTSIDE the tag; only the language tag and the verbatim, untranslated words sit inside. Speaker IDs (S1, S2) stay stable across shots. Speech truncated by the clip end gets <cutoff>. Voiceover requires the exact phrase "says in an off-screen voiceover" plus an immediate "lips remain closed" clause or the character lip-flaps.
- Stable dialogue languages: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish. Turkish is OUTSIDE the set: never invent a language tag — write the shot with no <d> block, describe the performance physically ("she speaks two short sentences to camera, lips moving with measured phrasing, no audible dialogue directed to the microphone"), set the room in overall_soundscape, and let the user lay ADR in post.
- Proper nouns BLOCK the request (character names, film titles, studios, celebrities trip input moderation). Describe the silhouette; a director's or auteur's STYLE name is safe ("the Hitchcock camera movement").
- "Cinematic" is legal exactly once, in the style slot at the head of [Shot 1] ("[Shot 1] Cinematic, live-action, a medium-wide shot frames …"); elsewhere it and "beautiful" are banned abstractions. Styles that work: Cinematic, live-action, 2D-animated, 3D CG, claymation, watercolor, vintage film.
- On-screen text goes in ASCII double quotes, verbatim, untranslated: A red neon sign reading "OPEN" glows above the doorway. Unquoted text returns as letter-shaped noise; one single-line caption on screen at a time.
- No symbols in the prose — no arrows, slashes, plus signs or connectors; the model paints them into the frame. No weights, no (word:1.3), no BREAK, no tag-soup.
- Exclusions are positive end-states first, and a short trailing "No X, no Y" clause only as backstop (MiniMax's own templates end "no split screen, no hard cuts, no random camera shake, no watermark").
- Keep faces medium or closer; a wide frontal face degrades regardless of resolution (vendor-acknowledged). Wide frames get back or rear-three-quarter views. Max three characters with on-screen action or dialogue per shot.
What works best
- Target 350-500 words of body; hard ceiling 7,000 characters.
- Two shots by default, three when the beats demand it; spend a cut only for new information (subject, space, state, viewpoint, time) — a change of distance is a camera move, not a cut. Beat budget: 5 s → 3-4 beats · 10 s → 5-7 · 15 s → 6-9, one primary action per beat.
- Dialogue density: a sentence or two per five seconds. Over-write and the delivery accelerates; under-write and the model fills the gap with invented speech. Two lines of dialogue in 15 s is the practical ceiling before background music silently vanishes.
- Diegetic sound (radio, busker, phone speaker, on-camera band) is written in the body next to the picture that motivates it; score goes in non_diegetic_music — a routing rule, not a style choice.
- Fast action shatters limbs and clothing: one action per shot, "with small amplitude at slow speed", and cut rather than accelerate.
- Timestamps address picture reliably and sound unreliably: schedule cuts, never a music hit against a frame.
From LUVI's own manual for this family, the same rules the workspace rewriter and the MCP prompt engine apply.