de xAIVideoImagen

Grok Imagine

12 modelos · desde 40 créditos

Modelos de la familia Grok Imagine

ModelModoPrecioRelaciones de aspecto
Grok Imagine Image 2.0 EditImage editdesde 140 créditos1:1 · 3:4 · 4:3 · 9:16 · 16:9 · 2:3Abrir
Grok Imagine Image 2.0 Text-to-ImageText to imagedesde 120 créditos1:1 · 3:4 · 4:3 · 9:16 · 16:9 · 2:3Abrir
Grok Imagine Image EditImage editdesde 44 créditos1:1 · 3:4 · 4:3 · 9:16 · 16:9 · 2:3Abrir
Grok Imagine Image-to-VideoImage to videodesde 1124 créditos1:1 · 16:9 · 9:16 · 4:3 · 3:4 · 3:2Abrir
Grok Imagine Quality EditImage editdesde 120 créditos1:1 · 3:4 · 4:3 · 9:16 · 16:9 · 2:3Abrir
Grok Imagine Quality Text-to-ImageText to imagedesde 100 créditos1:1 · 3:4 · 4:3 · 9:16 · 16:9 · 2:3Abrir
Grok Imagine Reference-to-VideoReference to videodesde 1124 créditos1:1 · 16:9 · 9:16 · 4:3 · 3:4 · 3:2Abrir
Grok Imagine Text-to-ImageText to imagedesde 40 créditos1:1 · 3:4 · 4:3 · 9:16 · 16:9 · 2:3Abrir
Grok Imagine Text-to-VideoText to videodesde 1120 créditos1:1 · 16:9 · 9:16 · 4:3 · 3:4 · 3:2Abrir
Grok Imagine Video EditVideo editdesde 1392 créditosAbrir
Grok Imagine Video ExtendVideo extenddesde 1140 créditosAbrir
Grok Imagine Video v1.5 Image-to-VideoImage to videodesde 2260 créditos1:1 · 16:9 · 9:16 · 4:3 · 3:4 · 3:2Abrir

Guía de prompts para Grok Imagine

TARGET MODEL: Grok Imagine (engine: Aurora — xAI autoregressive mixture-of-experts, NOT diffusion). It generates left-to-right like an LLM: THE FIRST SENTENCE CARRIES THE MOST WEIGHT.

Reglas que deben cumplirse

  • Flowing natural prose, never tag stacks. Order: subject + main action in the opening 20-30 words, then environment, then lighting, then style, then camera/technical.
  • Delete booster tags entirely: masterpiece, 8K, ultra detailed, ultra-realistic, breathtaking, stunning, hyperdetailed, best quality, award-winning. Delete weighting syntax: (term:1.4), (subject)++, [x], {x} — they do nothing on Aurora.
  • No negations: on an LLM-native model "no X" reintroduces X. Rephrase every exclusion as the desired positive state: "no blemishes" → "clear skin"; "no blur" → "sharp focus"; "no waves" → "calm sea"; "no people" → "empty, still scene". Two allowed idioms only: "no digital perfection, no smoothing" (anti-plastic skin) and "no additional text or labels" (suppress invented text).
  • Length 30-80 words. Cap the environment at 2-3 sensory elements — Aurora muddies busy scenes. Better direction beats more words.
  • In-image text: put the exact copy in quotes after a verb ("a sign reading 'OPEN LATE'"), short and single-line; font/color/effect as separate clauses; name the surface and placement ("neon on brick wall, top banner"). For non-Latin text write the actual glyphs. When the image should carry only that text, append "no additional text or labels".

Lo que mejor funciona

  • Photorealism: name the instrument, not the adjective. A camera brief beats the word "photorealistic": "Shot on Leica M10, 35mm Summilux, f/1.4, shallow depth of field, natural film grain". Bodies that land: Leica M10, Sony A7R V, Canon EOS R5, Fujifilm X-T4. Focal feel: 24mm establishing / 35mm street / 50mm natural / 85mm portrait. Film stock carries color+grain in one token: "Kodak Portra 400 color tones" (adds warmth against Grok's cool default), "35mm film grain".
  • Name lighting classically and date it: golden hour, Rembrandt, rim light, chiaroscuro, three-point, backlit, volumetric god rays. "Early March morning" beats "morning"; "overcast afternoon in November" beats "cloudy".
  • Name colors ("electric blue and hot pink", never "colorful"). Mood words only when evocative: nostalgic, melancholic, tense, dreamlike — never happy/nice/cool.
  • Reinforce one critical surface by restating it in different words ("chrome bodysuit; reflective chrome") — Aurora's substitute for weighting.
  • Style: at most ONE movement + ONE technique ("cyberpunk aesthetic with impressionist painting technique"). Grok defaults to glossy 3D and cool editorial tones — for flat mediums (vector, pixel art, watercolor, anime) name the medium explicitly and EARLY; anime needs a hard cue ("80s OVA anime, cel-shaded" or "Makoto Shinkai art style").
  • Physical-detail realism: skin pores, fabric weave, condensation, water droplets; add "candid, not posing" if the subject reads staged.
  • If the user stated an aspect ratio, optionally echo it as the sentence tail ("…documentary feel, 16:9").

Del manual propio de LUVI para esta familia: las mismas reglas que aplican el reescritor del espacio de trabajo y el motor de prompts de MCP.

Guía de prompts para Grok Imagine Image 2.0

TARGET MODEL: Grok Imagine Image 2.0 (xAI — Aurora lineage, an autoregressive mixture-of-experts network; NOT FLUX, never write or imply FLUX heritage). Its stated behaviour is that it "follows instructions closely, down to the details" and "plans typography and layout the way a designer would". Treat the prompt as a SPECIFICATION, not a mood board: write a brief.

Reglas que deben cumplirse

  • One flowing descriptive sentence or a small set of them, never comma-separated tag stacks. Every prompt xAI publishes is prose; commas do grammatical work (a head noun plus qualifying phrases), never enumeration. Test: if it cannot be read aloud as English, rewrite it.
  • Clause order — subject, then composition, then typography, then style/medium, then lighting, then camera, then material/finish, then palette/mood. Lead with the subject and the one or two attributes that make it THIS one; the layout planner needs a primary element before it can build hierarchy. Style/medium is the single highest-leverage phrase in the prompt — one noun there overrides a paragraph of adjectives.
  • Delete quality tokens outright: masterpiece, 8k, best quality, highly detailed, trending on artstation. They appear zero times in the entire official corpus. Delete weighting syntax too — (word:1.3), [word], :: — none of it is parsed and none of it exists in the API.
  • No negative prompt exists, and no seed exists. Convert every exclusion into the presence of its opposite: "no blemishes" becomes "smooth, clear skin"; "not blurry" becomes "sharp focus throughout, crisp edge detail"; "no text" becomes "a clean unlettered surface"; "no people" becomes "an empty street at dawn"; "not cluttered" becomes "generous negative space, three elements only"; "no modern objects" becomes "period-accurate 1890s dressing throughout". A negation clause is still tokens describing the forbidden thing.
  • Do NOT target a word count. xAI publishes no length guidance and its own prompts run 4-24 words; the community "30-80 word sweet spot" has no published methodology. Stop when every element you would reject the image for is specified, and not one clause later. Anything you would accept either way, leave out.
  • Cut competing constraints rather than adding adjectives. Close instruction-following averages incompatible requirements into blandness — a "wide establishing shot" of a "macro texture detail" returns mud. If two constraints cannot both be true, keep the one the user would reject the image over.
  • Aspect ratio, resolution, image count, output format and storage are parameters — never write them into the prompt. Do not write seed, cfg_scale, guidance_scale, steps, sampler or negative_prompt into the prose either; none of them exist on this model.
  • Specify the mechanism, not the adjective, wherever the user wants a specific result: "single hard key light from camera-left at 45 degrees, deep falloff, no fill" is an instruction; "dramatic lighting" is a judgement the model will answer with its own taste. Leave a slot underspecified only when the model solving it is genuinely wanted.

Lo que mejor funciona

  • Typography is this model's headline capability, and the artefact type does most of the work — "concert poster", "cereal box", "app onboarding screen", "airline safety card" each hand the planner a whole layout convention for free. Name the artefact, then the typographic hierarchy, then the legibility constraint. xAI's own flagship example is exactly that shape: "A concert poster for a synthwave band, bold retro typography, sharp small print".
  • Put in-image copy in quotation marks as a literal — the headline reads "MIDNIGHT DRIVE" — rather than describing it. Capitalising the literal is the most-cited field technique for making the model treat it as content to render rather than description; the side effect is set-in-caps type, so quote without capitals when mixed case matters.
  • Give the planner HIERARCHY, not a layout: "a large headline, a subhead half its size, and a four-line credit block at the base" is actionable. "well-balanced typography" is not.
  • Name the medium explicitly — photograph, illustration, oil, pencil, pixel art, stencil, woodblock, infographic. Camera language (lens, depth of field, angle) is only meaningful once the medium is photographic; on an illustration it is noise.
  • Lighting as direction plus quality plus source plus time of day, and date it where it matters: "soft window light from the left", "single soft key light from upper right, gentle reflection beneath".
  • Name colours rather than calling something colourful, and let palette be the closing modifier — it is the layer most tolerant of being left to the model.
  • Every distinct in-image string is another thing that must be spelled, positioned and kept legible at once. There is no published density threshold, so write the text the piece actually needs rather than pre-emptively simplifying.

Del manual propio de LUVI para esta familia: las mismas reglas que aplican el reescritor del espacio de trabajo y el motor de prompts de MCP.

Guía de prompts para Grok Imagine Video

TARGET MODEL: Grok Imagine Video (xAI Aurora — autoregressive; picture AND audio co-generate in the same pass). The prose is the whole instrument: no seed, no negative_prompt, no audio toggle.

Reglas que deben cumplirse

  • The formula, in timeline order: Subject + Action/Motion + Camera + Environment/Lighting + Style + Audio. Front-load the critical action — earlier words weigh more. 30-60 words optimal.
  • AUDIO IS ALWAYS ON with no silence flag: if the prompt doesn't specify sound, the clip gets silent or random audio — so ALWAYS write the audio in. Music, SFX timed to motion, ambient/room tone, lip-synced dialogue.
  • Dialogue: quote the line with a delivery-cue prefix — a quiet whisper: "We made it." · urgent shout: "Stop him!" — and keep lines SHORT (audio is the weakest layer; long lines turn to gibberish). Optional sound-design clause: AUDIO: soft room tone, faint kettle hiss.
  • ONE primary action and ONE named camera move per clip — conflicting moves (zoom+pan) break temporal coherence.
  • State motion magnitude explicitly — the model cannot infer it ("passing" → "passing quickly"). Strong verbs with intensity adverbs ("sprints", "surges", "pitches forward with tremendous force").
  • Negation is ignored — phrase every exclusion as the desired positive state ("sharp focus throughout", never "no blur").
  • One aesthetic per clip — never mix (no anime+photoreal). Lighting by source + direction.

Lo que mejor funciona

  • Camera vocabulary that lands: locked/static (strong default) · slow push-in · slow dolly-in · tracking shot alongside · handheld follow from behind · orbit/360 · slow crane pullback · rack focus · aerial push-in · pull-back · time-lapse. "Cinematic" or "dynamic camera" name nothing executable — name the move or lock the frame.
  • Avoid known failure modes: big body dynamics (jerky), extreme close-ups and macro hand work (distort), long photoreal holds (waxy morph). Favor short beats and implied/reaction motion.
  • Craft pattern for scenes with speech: scene → dialogue in quotes with its delivery cue → optional AUDIO: line.

Del manual propio de LUVI para esta familia: las mismas reglas que aplican el reescritor del espacio de trabajo y el motor de prompts de MCP.