by AlibabaImage

Qwen Image

10 models · from 42 credits

Models in the Qwen Image family

ModelModePriceAspect ratios
Qwen Image (Atlas) EditImage editfrom 64 creditsOpen
Qwen Image (Atlas) Text-to-ImageText to imagefrom 49 creditsOpen
Qwen Image 2.0 EditImage editfrom 56 creditsOpen
Qwen Image 2.0 Pro EditImage editfrom 120 creditsOpen
Qwen Image 2.0 Pro Text-to-ImageText to imagefrom 120 creditsOpen
Qwen Image 2.0 Text-to-ImageText to imagefrom 56 creditsOpen
Qwen Image EditImage editfrom 64 creditsOpen
Qwen Image Edit PlusImage editfrom 42 creditsOpen
Qwen Image MaxText to imagefrom 105 creditsOpen
Qwen Image PlusText to imagefrom 42 creditsOpen

Prompting guide for Qwen-Image

TARGET MODEL: Qwen-Image (Alibaba — text-to-image and instruction-driven image editing, incl. 2.0 / Plus / Max / Edit Plus tiers). The model's headline strength is rendering text — multi-line, paragraph-level, bilingual Chinese/English — correctly spelled and positioned, so the prompt is a structured description in which every visible string is quoted. Some rows expose a negative_prompt parameter; the positive prompt itself carries NO negative words.

Rules that must hold

  • Every piece of text that must appear in the image is enclosed in double quotes and given a position and treatment: 'a dark blue sign reads "Il Messaggero" in white Gothic lettering, top-left'. Keep the quoted words exactly as the user wrote them — same language, same capitalization — and never add a word of text the user did not ask for; the model paints whatever is quoted.
  • Structure the prompt as Subject + Setting + Style + Camera (shot size, angle, lens, composition) + Atmosphere + Detail modifiers, written as complete descriptive sentences — spatial relationships and shot composition are read literally ("center-right, a young woman …; below, a newsstand …").
  • No negative words in the prompt body ("no watermark", "without people" are not read; the negative_prompt parameter, where the row has one, is the place for "blur, extra fingers, garbled text").
  • Keep the rewritten prompt under ~200 words for a single image; the budget is 800 tokens on the async tiers and 1,300 on 2.0 — dense, not long.
  • Write in Chinese or English only (the supported languages); a non-English string that must render stays quoted, untranslated.
  • Choose one precise, named style when the user named none (commercial photography, watercolor, clay, ink painting, 3D cartoon, Pixar style, felt, origami, surrealism, pointillism) rather than a vague "artistic".

What works best

  • The vocabulary the guide itself uses: shot size (extreme close-up, close-up, medium shot, long shot) · perspective (eye level, bird's eye, low angle, aerial) · lens (macro, ultra-wide, telephoto, fisheye) · lighting (natural light, backlight, neon, ambient light, cinematic lighting, golden rim light) · composition (centered composition, half-body close-up).
  • For posters, UI, infographics and comics, describe the layout as a designer would — regions (top / center / bottom band), hierarchy (headline, subhead, caption), type treatment (bold sans-serif, handwritten, Gothic), color per element — and quote each string in its slot. Fewer simultaneous strings render cleaner.
  • For portraits state age, face shape, gaze, outfit, makeup and the light on the face; the official examples are full sentences of concrete visual fact, never mood words alone.
  • Quality tokens the official rewriter appends are fine at the tail, once: "Ultra HD, 4K, cinematic composition".

From LUVI's own manual for this family, the same rules the workspace rewriter and the MCP prompt engine apply.