Grok Imagine·by xAIImage to video

Grok Imagine Image-to-Video

xAI Grok Imagine Video — animates a starting frame with natural-language motion prompts. 1-15 seconds, 480p/720p.

Open in workspacexai/grok-imagine-video/image-to-video

Parameters

ParameterTypeDefaultRange or options
duration
duration

Video duration in seconds (1-15).

slider81 – 15
aspect ratio
aspect_ratio

Aspect ratio of the generated video.

select16:91:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3
resolution
resolution

Output resolution. 720p costs more per second than 480p.

select720p480p, 720p

Inputs

Files this model takes in addition to the prompt.

InputAcceptsMax files
image_url
image_url
image1Required

Prompting guide for Grok Imagine

TARGET MODEL: Grok Imagine Video (xAI Aurora — autoregressive; picture AND audio co-generate in the same pass). The prose is the whole instrument: no seed, no negative_prompt, no audio toggle.

Rules that must hold

  • The formula, in timeline order: Subject + Action/Motion + Camera + Environment/Lighting + Style + Audio. Front-load the critical action — earlier words weigh more. 30-60 words optimal.
  • AUDIO IS ALWAYS ON with no silence flag: if the prompt doesn't specify sound, the clip gets silent or random audio — so ALWAYS write the audio in. Music, SFX timed to motion, ambient/room tone, lip-synced dialogue.
  • Dialogue: quote the line with a delivery-cue prefix — a quiet whisper: "We made it." · urgent shout: "Stop him!" — and keep lines SHORT (audio is the weakest layer; long lines turn to gibberish). Optional sound-design clause: AUDIO: soft room tone, faint kettle hiss.
  • ONE primary action and ONE named camera move per clip — conflicting moves (zoom+pan) break temporal coherence.
  • State motion magnitude explicitly — the model cannot infer it ("passing" → "passing quickly"). Strong verbs with intensity adverbs ("sprints", "surges", "pitches forward with tremendous force").
  • Negation is ignored — phrase every exclusion as the desired positive state ("sharp focus throughout", never "no blur").
  • One aesthetic per clip — never mix (no anime+photoreal). Lighting by source + direction.

What works best

  • Camera vocabulary that lands: locked/static (strong default) · slow push-in · slow dolly-in · tracking shot alongside · handheld follow from behind · orbit/360 · slow crane pullback · rack focus · aerial push-in · pull-back · time-lapse. "Cinematic" or "dynamic camera" name nothing executable — name the move or lock the frame.
  • Avoid known failure modes: big body dynamics (jerky), extreme close-ups and macro hand work (distort), long photoreal holds (waxy morph). Favor short beats and implied/reaction motion.
  • Craft pattern for scenes with speech: scene → dialogue in quotes with its delivery cue → optional AUDIO: line.

From LUVI's own manual for this family, the same rules the workspace rewriter and the MCP prompt engine apply.

Other models in this family