von xAIVideoBild

Grok Imagine

12 Modelle · ab 40 Credits

Modelle der Familie Grok Imagine

ModelModusPreisSeitenverhältnisse
Grok Imagine Image 2.0 EditImage editab 140 Credits1:1 · 3:4 · 4:3 · 9:16 · 16:9 · 2:3Öffnen
Grok Imagine Image 2.0 Text-to-ImageText to imageab 120 Credits1:1 · 3:4 · 4:3 · 9:16 · 16:9 · 2:3Öffnen
Grok Imagine Image EditImage editab 44 Credits1:1 · 3:4 · 4:3 · 9:16 · 16:9 · 2:3Öffnen
Grok Imagine Image-to-VideoImage to videoab 1.124 Credits1:1 · 16:9 · 9:16 · 4:3 · 3:4 · 3:2Öffnen
Grok Imagine Quality EditImage editab 120 Credits1:1 · 3:4 · 4:3 · 9:16 · 16:9 · 2:3Öffnen
Grok Imagine Quality Text-to-ImageText to imageab 100 Credits1:1 · 3:4 · 4:3 · 9:16 · 16:9 · 2:3Öffnen
Grok Imagine Reference-to-VideoReference to videoab 1.124 Credits1:1 · 16:9 · 9:16 · 4:3 · 3:4 · 3:2Öffnen
Grok Imagine Text-to-ImageText to imageab 40 Credits1:1 · 3:4 · 4:3 · 9:16 · 16:9 · 2:3Öffnen
Grok Imagine Text-to-VideoText to videoab 1.120 Credits1:1 · 16:9 · 9:16 · 4:3 · 3:4 · 3:2Öffnen
Grok Imagine Video EditVideo editab 1.392 CreditsÖffnen
Grok Imagine Video ExtendVideo extendab 1.140 CreditsÖffnen
Grok Imagine Video v1.5 Image-to-VideoImage to videoab 2.260 Credits1:1 · 16:9 · 9:16 · 4:3 · 3:4 · 3:2Öffnen

Prompt-Leitfaden für Grok Imagine

TARGET MODEL: Grok Imagine (engine: Aurora — xAI autoregressive mixture-of-experts, NOT diffusion). It generates left-to-right like an LLM: THE FIRST SENTENCE CARRIES THE MOST WEIGHT.

Regeln, die gelten müssen

  • Flowing natural prose, never tag stacks. Order: subject + main action in the opening 20-30 words, then environment, then lighting, then style, then camera/technical.
  • Delete booster tags entirely: masterpiece, 8K, ultra detailed, ultra-realistic, breathtaking, stunning, hyperdetailed, best quality, award-winning. Delete weighting syntax: (term:1.4), (subject)++, [x], {x} — they do nothing on Aurora.
  • No negations: on an LLM-native model "no X" reintroduces X. Rephrase every exclusion as the desired positive state: "no blemishes" → "clear skin"; "no blur" → "sharp focus"; "no waves" → "calm sea"; "no people" → "empty, still scene". Two allowed idioms only: "no digital perfection, no smoothing" (anti-plastic skin) and "no additional text or labels" (suppress invented text).
  • Length 30-80 words. Cap the environment at 2-3 sensory elements — Aurora muddies busy scenes. Better direction beats more words.
  • In-image text: put the exact copy in quotes after a verb ("a sign reading 'OPEN LATE'"), short and single-line; font/color/effect as separate clauses; name the surface and placement ("neon on brick wall, top banner"). For non-Latin text write the actual glyphs. When the image should carry only that text, append "no additional text or labels".

Was am besten funktioniert

  • Photorealism: name the instrument, not the adjective. A camera brief beats the word "photorealistic": "Shot on Leica M10, 35mm Summilux, f/1.4, shallow depth of field, natural film grain". Bodies that land: Leica M10, Sony A7R V, Canon EOS R5, Fujifilm X-T4. Focal feel: 24mm establishing / 35mm street / 50mm natural / 85mm portrait. Film stock carries color+grain in one token: "Kodak Portra 400 color tones" (adds warmth against Grok's cool default), "35mm film grain".
  • Name lighting classically and date it: golden hour, Rembrandt, rim light, chiaroscuro, three-point, backlit, volumetric god rays. "Early March morning" beats "morning"; "overcast afternoon in November" beats "cloudy".
  • Name colors ("electric blue and hot pink", never "colorful"). Mood words only when evocative: nostalgic, melancholic, tense, dreamlike — never happy/nice/cool.
  • Reinforce one critical surface by restating it in different words ("chrome bodysuit; reflective chrome") — Aurora's substitute for weighting.
  • Style: at most ONE movement + ONE technique ("cyberpunk aesthetic with impressionist painting technique"). Grok defaults to glossy 3D and cool editorial tones — for flat mediums (vector, pixel art, watercolor, anime) name the medium explicitly and EARLY; anime needs a hard cue ("80s OVA anime, cel-shaded" or "Makoto Shinkai art style").
  • Physical-detail realism: skin pores, fabric weave, condensation, water droplets; add "candid, not posing" if the subject reads staged.
  • If the user stated an aspect ratio, optionally echo it as the sentence tail ("…documentary feel, 16:9").

Aus LUVIs eigenem Handbuch für diese Familie; dieselben Regeln wendet der Rewriter im Arbeitsbereich und die MCP-Prompt-Engine an.

Prompt-Leitfaden für Grok Imagine Image 2.0

TARGET MODEL: Grok Imagine Image 2.0 (xAI — Aurora lineage, an autoregressive mixture-of-experts network; NOT FLUX, never write or imply FLUX heritage). Its stated behaviour is that it "follows instructions closely, down to the details" and "plans typography and layout the way a designer would". Treat the prompt as a SPECIFICATION, not a mood board: write a brief.

Regeln, die gelten müssen

  • One flowing descriptive sentence or a small set of them, never comma-separated tag stacks. Every prompt xAI publishes is prose; commas do grammatical work (a head noun plus qualifying phrases), never enumeration. Test: if it cannot be read aloud as English, rewrite it.
  • Clause order — subject, then composition, then typography, then style/medium, then lighting, then camera, then material/finish, then palette/mood. Lead with the subject and the one or two attributes that make it THIS one; the layout planner needs a primary element before it can build hierarchy. Style/medium is the single highest-leverage phrase in the prompt — one noun there overrides a paragraph of adjectives.
  • Delete quality tokens outright: masterpiece, 8k, best quality, highly detailed, trending on artstation. They appear zero times in the entire official corpus. Delete weighting syntax too — (word:1.3), [word], :: — none of it is parsed and none of it exists in the API.
  • No negative prompt exists, and no seed exists. Convert every exclusion into the presence of its opposite: "no blemishes" becomes "smooth, clear skin"; "not blurry" becomes "sharp focus throughout, crisp edge detail"; "no text" becomes "a clean unlettered surface"; "no people" becomes "an empty street at dawn"; "not cluttered" becomes "generous negative space, three elements only"; "no modern objects" becomes "period-accurate 1890s dressing throughout". A negation clause is still tokens describing the forbidden thing.
  • Do NOT target a word count. xAI publishes no length guidance and its own prompts run 4-24 words; the community "30-80 word sweet spot" has no published methodology. Stop when every element you would reject the image for is specified, and not one clause later. Anything you would accept either way, leave out.
  • Cut competing constraints rather than adding adjectives. Close instruction-following averages incompatible requirements into blandness — a "wide establishing shot" of a "macro texture detail" returns mud. If two constraints cannot both be true, keep the one the user would reject the image over.
  • Aspect ratio, resolution, image count, output format and storage are parameters — never write them into the prompt. Do not write seed, cfg_scale, guidance_scale, steps, sampler or negative_prompt into the prose either; none of them exist on this model.
  • Specify the mechanism, not the adjective, wherever the user wants a specific result: "single hard key light from camera-left at 45 degrees, deep falloff, no fill" is an instruction; "dramatic lighting" is a judgement the model will answer with its own taste. Leave a slot underspecified only when the model solving it is genuinely wanted.

Was am besten funktioniert

  • Typography is this model's headline capability, and the artefact type does most of the work — "concert poster", "cereal box", "app onboarding screen", "airline safety card" each hand the planner a whole layout convention for free. Name the artefact, then the typographic hierarchy, then the legibility constraint. xAI's own flagship example is exactly that shape: "A concert poster for a synthwave band, bold retro typography, sharp small print".
  • Put in-image copy in quotation marks as a literal — the headline reads "MIDNIGHT DRIVE" — rather than describing it. Capitalising the literal is the most-cited field technique for making the model treat it as content to render rather than description; the side effect is set-in-caps type, so quote without capitals when mixed case matters.
  • Give the planner HIERARCHY, not a layout: "a large headline, a subhead half its size, and a four-line credit block at the base" is actionable. "well-balanced typography" is not.
  • Name the medium explicitly — photograph, illustration, oil, pencil, pixel art, stencil, woodblock, infographic. Camera language (lens, depth of field, angle) is only meaningful once the medium is photographic; on an illustration it is noise.
  • Lighting as direction plus quality plus source plus time of day, and date it where it matters: "soft window light from the left", "single soft key light from upper right, gentle reflection beneath".
  • Name colours rather than calling something colourful, and let palette be the closing modifier — it is the layer most tolerant of being left to the model.
  • Every distinct in-image string is another thing that must be spelled, positioned and kept legible at once. There is no published density threshold, so write the text the piece actually needs rather than pre-emptively simplifying.

Aus LUVIs eigenem Handbuch für diese Familie; dieselben Regeln wendet der Rewriter im Arbeitsbereich und die MCP-Prompt-Engine an.

Prompt-Leitfaden für Grok Imagine Video

TARGET MODEL: Grok Imagine Video (xAI Aurora — autoregressive; picture AND audio co-generate in the same pass). The prose is the whole instrument: no seed, no negative_prompt, no audio toggle.

Regeln, die gelten müssen

  • The formula, in timeline order: Subject + Action/Motion + Camera + Environment/Lighting + Style + Audio. Front-load the critical action — earlier words weigh more. 30-60 words optimal.
  • AUDIO IS ALWAYS ON with no silence flag: if the prompt doesn't specify sound, the clip gets silent or random audio — so ALWAYS write the audio in. Music, SFX timed to motion, ambient/room tone, lip-synced dialogue.
  • Dialogue: quote the line with a delivery-cue prefix — a quiet whisper: "We made it." · urgent shout: "Stop him!" — and keep lines SHORT (audio is the weakest layer; long lines turn to gibberish). Optional sound-design clause: AUDIO: soft room tone, faint kettle hiss.
  • ONE primary action and ONE named camera move per clip — conflicting moves (zoom+pan) break temporal coherence.
  • State motion magnitude explicitly — the model cannot infer it ("passing" → "passing quickly"). Strong verbs with intensity adverbs ("sprints", "surges", "pitches forward with tremendous force").
  • Negation is ignored — phrase every exclusion as the desired positive state ("sharp focus throughout", never "no blur").
  • One aesthetic per clip — never mix (no anime+photoreal). Lighting by source + direction.

Was am besten funktioniert

  • Camera vocabulary that lands: locked/static (strong default) · slow push-in · slow dolly-in · tracking shot alongside · handheld follow from behind · orbit/360 · slow crane pullback · rack focus · aerial push-in · pull-back · time-lapse. "Cinematic" or "dynamic camera" name nothing executable — name the move or lock the frame.
  • Avoid known failure modes: big body dynamics (jerky), extreme close-ups and macro hand work (distort), long photoreal holds (waxy morph). Favor short beats and implied/reaction motion.
  • Craft pattern for scenes with speech: scene → dialogue in quotes with its delivery cue → optional AUDIO: line.

Aus LUVIs eigenem Handbuch für diese Familie; dieselben Regeln wendet der Rewriter im Arbeitsbereich und die MCP-Prompt-Engine an.