Grok Imagine·by xAIImage edit
Grok Imagine Image 2.0 Edit
Natural-language image editing with xAI Grok Imagine Image 2.0. Up to 3 source images (+$0.01 each); multi-image editing supports <IMAGE_0>, <IMAGE_1> notation. Output aspect ratio follows the first source image.
xai/grok-imagine-image-2.0/editParameters
| Parameter | Type | Default | Range or options |
|---|---|---|---|
aspect ratio aspect_ratioAspect ratio of the edited image. auto follows the first source image. | select | auto | auto, 1:1, 3:4, 4:3, 9:16, 16:9, 2:3, 3:2, 9:19.5, 19.5:9, 9:20, 20:9, 1:2, 2:1 |
Resolution resolutionOutput resolution. 1k = 1024×1024, 2k = 2048×2048 (+$0.02 per image). | select | 1k | 1k, 2k |
num images num_imagesNumber of images to generate. Each image is billed separately. | select | 1 | 1, 2, 3, 4 |
Quality qualityRendering quality tier. low is cheaper and roughly 8× faster; medium is the default and produces the model's best output. | select | medium | low, medium |
Inputs
Files this model takes in addition to the prompt.
| Input | Accepts | Max files | |
|---|---|---|---|
image_urls image_urls | image | 3 | Required |
Prompting guide for Grok Imagine
TARGET MODEL: Grok Imagine Image 2.0 (xAI — Aurora lineage, an autoregressive mixture-of-experts network; NOT FLUX, never write or imply FLUX heritage). Its stated behaviour is that it "follows instructions closely, down to the details" and "plans typography and layout the way a designer would". Treat the prompt as a SPECIFICATION, not a mood board: write a brief.
Rules that must hold
- One flowing descriptive sentence or a small set of them, never comma-separated tag stacks. Every prompt xAI publishes is prose; commas do grammatical work (a head noun plus qualifying phrases), never enumeration. Test: if it cannot be read aloud as English, rewrite it.
- Clause order — subject, then composition, then typography, then style/medium, then lighting, then camera, then material/finish, then palette/mood. Lead with the subject and the one or two attributes that make it THIS one; the layout planner needs a primary element before it can build hierarchy. Style/medium is the single highest-leverage phrase in the prompt — one noun there overrides a paragraph of adjectives.
- Delete quality tokens outright: masterpiece, 8k, best quality, highly detailed, trending on artstation. They appear zero times in the entire official corpus. Delete weighting syntax too — (word:1.3), [word], :: — none of it is parsed and none of it exists in the API.
- No negative prompt exists, and no seed exists. Convert every exclusion into the presence of its opposite: "no blemishes" becomes "smooth, clear skin"; "not blurry" becomes "sharp focus throughout, crisp edge detail"; "no text" becomes "a clean unlettered surface"; "no people" becomes "an empty street at dawn"; "not cluttered" becomes "generous negative space, three elements only"; "no modern objects" becomes "period-accurate 1890s dressing throughout". A negation clause is still tokens describing the forbidden thing.
- Do NOT target a word count. xAI publishes no length guidance and its own prompts run 4-24 words; the community "30-80 word sweet spot" has no published methodology. Stop when every element you would reject the image for is specified, and not one clause later. Anything you would accept either way, leave out.
- Cut competing constraints rather than adding adjectives. Close instruction-following averages incompatible requirements into blandness — a "wide establishing shot" of a "macro texture detail" returns mud. If two constraints cannot both be true, keep the one the user would reject the image over.
- Aspect ratio, resolution, image count, output format and storage are parameters — never write them into the prompt. Do not write seed, cfg_scale, guidance_scale, steps, sampler or negative_prompt into the prose either; none of them exist on this model.
- Specify the mechanism, not the adjective, wherever the user wants a specific result: "single hard key light from camera-left at 45 degrees, deep falloff, no fill" is an instruction; "dramatic lighting" is a judgement the model will answer with its own taste. Leave a slot underspecified only when the model solving it is genuinely wanted.
What works best
- Typography is this model's headline capability, and the artefact type does most of the work — "concert poster", "cereal box", "app onboarding screen", "airline safety card" each hand the planner a whole layout convention for free. Name the artefact, then the typographic hierarchy, then the legibility constraint. xAI's own flagship example is exactly that shape: "A concert poster for a synthwave band, bold retro typography, sharp small print".
- Put in-image copy in quotation marks as a literal — the headline reads "MIDNIGHT DRIVE" — rather than describing it. Capitalising the literal is the most-cited field technique for making the model treat it as content to render rather than description; the side effect is set-in-caps type, so quote without capitals when mixed case matters.
- Give the planner HIERARCHY, not a layout: "a large headline, a subhead half its size, and a four-line credit block at the base" is actionable. "well-balanced typography" is not.
- Name the medium explicitly — photograph, illustration, oil, pencil, pixel art, stencil, woodblock, infographic. Camera language (lens, depth of field, angle) is only meaningful once the medium is photographic; on an illustration it is noise.
- Lighting as direction plus quality plus source plus time of day, and date it where it matters: "soft window light from the left", "single soft key light from upper right, gentle reflection beneath".
- Name colours rather than calling something colourful, and let palette be the closing modifier — it is the layer most tolerant of being left to the model.
- Every distinct in-image string is another thing that must be spelled, positioned and kept legible at once. There is no published density threshold, so write the text the piece actually needs rather than pre-emptively simplifying.
From LUVI's own manual for this family, the same rules the workspace rewriter and the MCP prompt engine apply.