MiniMax H3·MiniMaxReference to videoYeni

MiniMax H3 Fast Reference-to-Video

Referans görseller, videolar ve bir ses parçasıyla (en fazla 9 görsel, 3 video, 3 ses) özneleri tutarlı tutarak MiniMax H3 Fast ile video üretin — hızlı deneme için en hızlı ve en ekonomik 480P kademesi. Referans videolar kendi süresiyle ayrıca ücretlendirilir. 480P; 5-15 saniye.

En uygun olduğu işler

Karakter tutarlılığıÜrün yerleştirmeMüzik/ritim güdümlü videoÇoklu referanslı sahneler
Çalışma alanında açminimax/h3-fast/reference-to-video

Parametreler

ParametreTürVarsayılanAralık veya seçenekler
Süre (sn)
duration

Video süresi (saniye, 5-15).

select85, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15
Çözünürlük
resolution

Video çözünürlüğü. Saniye fiyatı: 480P $0.0437/s.

select480P480P (480P)
En-Boy Oranı
ratio

En-boy oranı. Fiyatı etkilemez.

select16:921:9, 16:9, 4:3, 1:1, 3:4, 9:16

Girdiler

Prompt'a ek olarak bu modelin aldığı dosyalar.

GirdiKabul ederEn fazla dosya
Referans Görseller
refers_images
image9İsteğe bağlı
Referans Videolar
refers_videos
video3İsteğe bağlı
Referans Sesler
refers_audios
audio3İsteğe bağlı

MiniMax H3 için prompt rehberi

TARGET MODEL: MiniMax H3 (V2 endpoint — H3 and H3-Max). Joint audio-video in one pass: voice, sound effects and music are modeled together and are ALWAYS generated; there is no mute flag, no negative_prompt, no weighting syntax, no seed. The prompt is a fielded document in H3's own grammar, not a caption — the rewriter's job is to emit that document.

Uyulması gereken kurallar

  • Emit the fielded structure, lowercase field names, colon-terminated, one blank line between fields: integrated_multimodal_description: … / overall_soundscape: … / non_diegetic_music: … — in that order, every time.
  • ALWAYS emit BOTH audio fields. An omitted field is filled with invented sound, never silence. overall_soundscape = 1-4 sentences in one paragraph: ambience, physical action sounds, non-verbal human sounds — no dialogue, no diegetic music. non_diegetic_music = 1-3 sentences on instrumentation, tempo, rhythm and dynamic change — never mood words, never the emotional function of the score. Silence is the literal N/A (music N/A is reliable; soundscape N/A only when the user asks for a fully silent clip).
  • Camera is PROSE, never a bracket. Twelve moves, exact spelling: Zoom In/Out · Push In/Pull Out · Pan Left/Right · Truck Left/Right · Tilt Up/Down · Pedestal Up/Down · Arc Shot · Tracking Shot · Static Shot · Shake Slightly/Strongly · POV · Roll Clockwise/Counterclockwise. Four modifier literals only: with small amplitude · with large amplitude · at slow speed · at fast speed (medium amplitude / normal speed are expressed by omission). Form: "The camera pushes in with small amplitude at slow speed toward the folded letter in her hands." Never [Push in], [pan], [static], "dolly in", ECU/MCU/OTS, or lens/f-stop numbers — write "shallow depth of field" and lowercase shot sizes ("a medium-wide shot frames…").
  • One shot, one move, one visible action, one end frame. A second camera move costs a cut, not a comma: "[Shot 2] At 00:03.500, the camera cuts to …" — no timestamp on Shot 1, strictly increasing MM:SS.mmm cut times inside the clip, cut verbs from the allowlist only (the camera cuts to · the shot cuts to · the shot transitions to · the shot changes to · the shot switches to). Cross-dissolves and fades only when the user asked; on a hard-cut clip carry a standing "No dissolves, no morph transitions" clause because the model drifts toward soft transitions.
  • Dialogue is a tagged block: The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d> — delivery, action and speaker ID sit OUTSIDE the tag; only the language tag and the verbatim, untranslated words sit inside. Speaker IDs (S1, S2) stay stable across shots. Speech truncated by the clip end gets <cutoff>. Voiceover requires the exact phrase "says in an off-screen voiceover" plus an immediate "lips remain closed" clause or the character lip-flaps.
  • Stable dialogue languages: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish. Turkish is OUTSIDE the set: never invent a language tag — write the shot with no <d> block, describe the performance physically ("she speaks two short sentences to camera, lips moving with measured phrasing, no audible dialogue directed to the microphone"), set the room in overall_soundscape, and let the user lay ADR in post.
  • Proper nouns BLOCK the request (character names, film titles, studios, celebrities trip input moderation). Describe the silhouette; a director's or auteur's STYLE name is safe ("the Hitchcock camera movement").
  • "Cinematic" is legal exactly once, in the style slot at the head of [Shot 1] ("[Shot 1] Cinematic, live-action, a medium-wide shot frames …"); elsewhere it and "beautiful" are banned abstractions. Styles that work: Cinematic, live-action, 2D-animated, 3D CG, claymation, watercolor, vintage film.
  • On-screen text goes in ASCII double quotes, verbatim, untranslated: A red neon sign reading "OPEN" glows above the doorway. Unquoted text returns as letter-shaped noise; one single-line caption on screen at a time.
  • No symbols in the prose — no arrows, slashes, plus signs or connectors; the model paints them into the frame. No weights, no (word:1.3), no BREAK, no tag-soup.
  • Exclusions are positive end-states first, and a short trailing "No X, no Y" clause only as backstop (MiniMax's own templates end "no split screen, no hard cuts, no random camera shake, no watermark").
  • Keep faces medium or closer; a wide frontal face degrades regardless of resolution (vendor-acknowledged). Wide frames get back or rear-three-quarter views. Max three characters with on-screen action or dialogue per shot.

En iyi sonuç verenler

  • Target 350-500 words of body; hard ceiling 7,000 characters.
  • Two shots by default, three when the beats demand it; spend a cut only for new information (subject, space, state, viewpoint, time) — a change of distance is a camera move, not a cut. Beat budget: 5 s → 3-4 beats · 10 s → 5-7 · 15 s → 6-9, one primary action per beat.
  • Dialogue density: a sentence or two per five seconds. Over-write and the delivery accelerates; under-write and the model fills the gap with invented speech. Two lines of dialogue in 15 s is the practical ceiling before background music silently vanishes.
  • Diegetic sound (radio, busker, phone speaker, on-camera band) is written in the body next to the picture that motivates it; score goes in non_diegetic_music — a routing rule, not a style choice.
  • Fast action shatters limbs and clothing: one action per shot, "with small amplitude at slow speed", and cut rather than accelerate.
  • Timestamps address picture reliably and sound unreliably: schedule cuts, never a music hit against a frame.

LUVI'nin bu aile için yazdığı manuelden; çalışma alanındaki yeniden yazıcı ve MCP prompt motoru aynı kuralları uygular.

Bu ailedeki diğer modeller