GPT Image·von OpenAIText to image

Openai GPT Image 2 Text-to-Image

GPT Image 2 text to image is OpenAI's fast, cost-efficient text-to-image generator powered by GPT-5 guidance. Create photorealistic shots, product renders, concept art, and stylized graphics from natural-language prompts (optionally conditioned with an image). Supports custom aspect ratios, seeds, negative prompts, hex color hints, and style presets. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Im Arbeitsbereich öffnenopenai/gpt-image-2/text-to-image

Parameter

ParameterTypStandardBereich oder Optionen
output format
output_format

The format of the output image.

selectjpegjpeg, png
moderation
moderation

Whether to enable content moderation. If enabled, the system will check the input prompt for potentially harmful content and reject requests that violate content policies. This property is only available through the API.

textarealow
quality
quality

The quality of the generated image.

selectmediumlow, medium, high
size
size

optional string or auto or 1024x1024 or 1536x1024 or 5 moreThe size ofthe generated images.For gpt-image-2 and gpt-image-2-2026-04-21,arbitraryresolutions are supported as WIDTHxHEIGHT strings,for example 1536x864 . Width and height mustboth be divisible by 16 and the requested aspect ratio must be between 1:3 and 3:1.Resolutions above2560x1440 are experimental, and the maximum supported resolution is 3840x2160 .The requesteosize must also satisfy the model's current pixel and edge limits.The standard sizes 1024x1024.1536x1024 ,and 1024x1536 are supported by the GPT image models; auto is supported for modelsthat allow automatic sizing.For dall-e-2,use oneof 256x256,512x512,or 1024x1024.Fordall-e-3,use oneof 1024x1024,1792x1024,or 1024x1792.

select1024x10241024x768, 768x1024, 1024x1024, 1024x1536, 1536x1024, 2560x1440, 1440x2560, 3840x2160, 2160x3840

Eingaben

Keine Referenzdateien: dieses Modell arbeitet allein mit dem Prompt.

Prompt-Leitfaden für GPT Image

TARGET MODEL: GPT Image 2 (OpenAI). Prose + reference images are the ENTIRE control surface — no seed, no negative_prompt field, no style dial.

Regeln, die gelten müssen

  • Prose, not tags. Natural flowing English; commas only for concrete fragments, never abstract quality stacks.
  • Ban dead tokens: never "8k, ultra-detailed, masterpiece, best quality, highly detailed, 1girl" — inert at best, degrading at worst (they feed the tiling artifact).
  • Order: scene/background → subject → key details → constraints. Constraints ALWAYS last. Lead with the subject only when the figure is the hero — constraints still close.
  • Name the deliverable in one word: "editorial photo", "product shot", "UI mock", "infographic", "poster" — it sets the model's mode and polish level.
  • Concrete and physical, not adjectival: "warm tungsten key from the left, shallow depth of field" beats "cinematic lighting". A lens is a feel cue only ("50mm feel") — the model runs no optical simulation.
  • Affirmative beats negative. Describe what IS, materially and confidently. Exclusions are SHORT concrete noun tokens in the constraints slot ("no watermark, no extra text") — never long "do not include…" sentences (they inject the banned concept).
  • Aspect ratio is a parameter, never a sentence — do not write "16:9" or framing ratios into the prose.
  • Never ask for a transparent background in prose — it paints a literal checkerboard as solid pixels.
  • Start lean — don't overload one brief; a shorter concrete prompt beats a maximal one.

Was am besten funktioniert

  • Candid/authentic register — describe the DEGRADATION, not the production: "photorealistic" + one of "candid photograph / amateur photograph / iPhone photo", then the stack: "amateur composition, no studio lighting, subtle film grain, slight sensor noise in darker areas, slight overexposed highlights, mild compression artifacts, imperfect framing, natural color balance". Close on subtraction: "No glamorization, no heavy retouching. No cinematic grading, no flash, no studio light." Beat the camera-flash default with practical sources + two mixed color temperatures. Kill the influencer face: "no beauty filter, no HD polish".
  • Commercial/product register is the OPPOSITE — keep studio language: "premium product photography, studio lighting, seamless white sweep, soft shadow beneath, sharp label printing". Pick the register from the stated use.
  • Tiling guard on foliage / fur / grass / dense organic texture: shorten the prompt, drop grain vocabulary, and add "no speckled dot artifacts, no tiling texture, no repetitive grime pattern, no stippled texture".
  • Style needs visual targets, not adjectives: "cream background, heavy black condensed sans, one hero object, generous negative space" beats "minimalist premium editorial". Name the medium flat ("35mm film photograph", "flat vector design", "hand-painted watercolor"); hyper-specific device/era mediums land ("early 2000s CCTV security camera", "low-poly PlayStation screengrab").
  • Fight the warm/yellow house cast with explicit "natural color balance".
  • In-image text: quote the literal copy in "double quotes" or ALL CAPS; describe type as a separate spatial constraint (weight, color, placement); spell tricky or brand words letter-by-letter; keep each text zone to 5 words or fewer and bind distinct strings to distinct objects ("Include ONLY this label text (verbatim): …"); multiple text blocks in one frame MERGE — keep blocks few; fence phantom text every time: "No other visible words. No random letters. No duplicate text. No watermark." and assert the single instance ("Render the tagline exactly once.").
  • Character consistency is prompt-only — restate the preserve-list every turn ("change only X, keep everything else the same") and enumerate the identity traits; pair PRESERVE attributes with EXCLUDE additions ("no text, no watermark, no logos"). One change per edit.

Aus LUVIs eigenem Handbuch für diese Familie; dieselben Regeln wendet der Rewriter im Arbeitsbereich und die MCP-Prompt-Engine an.

Weitere Modelle dieser Familie