Qwen Image·de AlibabaText to image
Qwen Image 2.0 Text-to-Image
Alibaba Qwen Image 2.0 — fast, capable text-to-image. Flexible size, flat price.
qwen/qwen-image-2.0/text-to-imageParámetros
| Parámetro | Tipo | Por defecto | Rango u opciones |
|---|---|---|---|
size sizeOutput image size in WIDTH*HEIGHT pixels. | textarea | 1024*1024 | — |
Entradas
Sin archivos de referencia: este modelo trabaja solo con el prompt.
Guía de prompts para Qwen Image
TARGET MODEL: Qwen-Image (Alibaba — text-to-image and instruction-driven image editing, incl. 2.0 / Plus / Max / Edit Plus tiers). The model's headline strength is rendering text — multi-line, paragraph-level, bilingual Chinese/English — correctly spelled and positioned, so the prompt is a structured description in which every visible string is quoted. Some rows expose a negative_prompt parameter; the positive prompt itself carries NO negative words.
Reglas que deben cumplirse
- Every piece of text that must appear in the image is enclosed in double quotes and given a position and treatment: 'a dark blue sign reads "Il Messaggero" in white Gothic lettering, top-left'. Keep the quoted words exactly as the user wrote them — same language, same capitalization — and never add a word of text the user did not ask for; the model paints whatever is quoted.
- Structure the prompt as Subject + Setting + Style + Camera (shot size, angle, lens, composition) + Atmosphere + Detail modifiers, written as complete descriptive sentences — spatial relationships and shot composition are read literally ("center-right, a young woman …; below, a newsstand …").
- No negative words in the prompt body ("no watermark", "without people" are not read; the negative_prompt parameter, where the row has one, is the place for "blur, extra fingers, garbled text").
- Keep the rewritten prompt under ~200 words for a single image; the budget is 800 tokens on the async tiers and 1,300 on 2.0 — dense, not long.
- Write in Chinese or English only (the supported languages); a non-English string that must render stays quoted, untranslated.
- Choose one precise, named style when the user named none (commercial photography, watercolor, clay, ink painting, 3D cartoon, Pixar style, felt, origami, surrealism, pointillism) rather than a vague "artistic".
Lo que mejor funciona
- The vocabulary the guide itself uses: shot size (extreme close-up, close-up, medium shot, long shot) · perspective (eye level, bird's eye, low angle, aerial) · lens (macro, ultra-wide, telephoto, fisheye) · lighting (natural light, backlight, neon, ambient light, cinematic lighting, golden rim light) · composition (centered composition, half-body close-up).
- For posters, UI, infographics and comics, describe the layout as a designer would — regions (top / center / bottom band), hierarchy (headline, subhead, caption), type treatment (bold sans-serif, handwritten, Gothic), color per element — and quote each string in its slot. Fewer simultaneous strings render cleaner.
- For portraits state age, face shape, gaze, outfit, makeup and the light on the face; the official examples are full sentences of concrete visual fact, never mood words alone.
- Quality tokens the official rewriter appends are fine at the tail, once: "Ultra HD, 4K, cinematic composition".
Del manual propio de LUVI para esta familia: las mismas reglas que aplican el reescritor del espacio de trabajo y el motor de prompts de MCP.