Grok Imagine·par xAIText to video
Grok Imagine Text-to-Video
xAI Grok Imagine Video — generates 1-15 second videos from natural-language prompts. 480p/720p, 7 aspect ratios.
Ouvrir dans l'espace de travail →
xai/grok-imagine-video/text-to-videoParamètres
| Paramètre | Type | Par défaut | Plage ou options |
|---|---|---|---|
duration durationVideo duration in seconds (1-15). | slider | 8 | 1 – 15 |
aspect ratio aspect_ratioAspect ratio of the generated video. | select | 16:9 | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3 |
resolution resolutionOutput resolution. 720p costs more per second than 480p. | select | 720p | 480p, 720p |
Entrées
Aucun fichier de référence : ce modèle travaille à partir du prompt seul.
Guide de prompt pour Grok Imagine
TARGET MODEL: Grok Imagine Video (xAI Aurora — autoregressive; picture AND audio co-generate in the same pass). The prose is the whole instrument: no seed, no negative_prompt, no audio toggle.
Règles à respecter
- The formula, in timeline order: Subject + Action/Motion + Camera + Environment/Lighting + Style + Audio. Front-load the critical action — earlier words weigh more. 30-60 words optimal.
- AUDIO IS ALWAYS ON with no silence flag: if the prompt doesn't specify sound, the clip gets silent or random audio — so ALWAYS write the audio in. Music, SFX timed to motion, ambient/room tone, lip-synced dialogue.
- Dialogue: quote the line with a delivery-cue prefix — a quiet whisper: "We made it." · urgent shout: "Stop him!" — and keep lines SHORT (audio is the weakest layer; long lines turn to gibberish). Optional sound-design clause: AUDIO: soft room tone, faint kettle hiss.
- ONE primary action and ONE named camera move per clip — conflicting moves (zoom+pan) break temporal coherence.
- State motion magnitude explicitly — the model cannot infer it ("passing" → "passing quickly"). Strong verbs with intensity adverbs ("sprints", "surges", "pitches forward with tremendous force").
- Negation is ignored — phrase every exclusion as the desired positive state ("sharp focus throughout", never "no blur").
- One aesthetic per clip — never mix (no anime+photoreal). Lighting by source + direction.
Ce qui fonctionne le mieux
- Camera vocabulary that lands: locked/static (strong default) · slow push-in · slow dolly-in · tracking shot alongside · handheld follow from behind · orbit/360 · slow crane pullback · rack focus · aerial push-in · pull-back · time-lapse. "Cinematic" or "dynamic camera" name nothing executable — name the move or lock the frame.
- Avoid known failure modes: big body dynamics (jerky), extreme close-ups and macro hand work (distort), long photoreal holds (waxy morph). Favor short beats and implied/reaction motion.
- Craft pattern for scenes with speech: scene → dialogue in quotes with its delivery cue → optional AUDIO: line.
Issu du manuel LUVI pour cette famille : les mêmes règles qu'appliquent le réécriveur de l'espace de travail et le moteur de prompt MCP.
Autres modèles de cette famille
Grok Imagine Image 2.0 EditGrok Imagine Image 2.0 Text-to-ImageGrok Imagine Image EditGrok Imagine Image-to-VideoGrok Imagine Quality EditGrok Imagine Quality Text-to-ImageGrok Imagine Reference-to-VideoGrok Imagine Text-to-ImageGrok Imagine Video EditGrok Imagine Video ExtendGrok Imagine Video v1.5 Image-to-Video