How do I pick a reference from my library in the Workspace?

Open the library drawer at the bottom of the Workspace and drag a file onto a reference slot, or right-click it and choose Use as Reference; you can also drop a file from your desktop or paste an image straight onto a slot.

Last updated:

The reference slots

When the selected model accepts input media, slots appear under the prompt — one per kind of input the model takes (image, video, audio, 3D), each with a label and a capacity. A slot marked required must be filled before Generate is enabled.

Four ways to fill a slot

  1. Library drawer. Press the library button in the Workspace; a drawer opens along the bottom with your uploads and generations, filtered and sorted like the main Library. Drag a tile onto a slot.
  2. Use as Reference. Right-click (or ⋮) a tile in the drawer and choose Use as Reference; it goes to the first slot that accepts its type.
  3. Drop from the desktop. Drag a file from your computer onto a slot or anywhere on the reference area; it uploads and fills the slot.
  4. Paste. Copy an image and paste it onto a slot.

On phones the reference row opens the same picker as a sheet, plus Upload from this phone.

What can be picked

Images, videos, audio files and 3D models — anything in your library that matches a slot's type. Transcripts and PDFs are not references and do not appear as options.

A result as a reference

Every result in the strip can be reused directly: right-click → Use as Reference. That is how an image becomes the first frame of a video, or a generated voice drives a lip-sync model.

Changing the model

Slots follow the model. Switching to a model with the same kind of slot keeps the reference; switching to one without that slot drops it.

Frequently asked

How many references can I add?

As many as the model's slots allow — each slot shows its capacity. A multi-image edit model may take ten; a lip-sync model takes one image and one audio.

Why can't I pick a transcript?

A speech-to-text result is text, and no model takes text as a media reference. Copy it from the transcript viewer instead.

Do references cost extra?

A few image models charge a small fee per reference image; audio- and video-driven models bill on the reference's length. The estimate includes it.

More questions