Skip to main content
Multi-reference editing combines several input images into one output. Use it to restyle a photo after an artwork, put a product into a scene, dress a person in garments from other photos, or keep a character consistent across variations. In the prompt, clarify the purpose of each image and how the pieces fit together.

FLUX 3 example: put an object from one image into another

Image 1 is the scene and image 2 supplies the object, and the prompt names both roles. It also says how the new object should fit: the van takes the lighting of image 1. The result is landscape like the turntable because aspect_ratio: "auto" follows the first reference.

How FLUX 3 reads multiple references

Send up to 10 images. FLUX 3 Image takes 1 to 10 entries in images, as URLs or base64 strings. Each must be at least 256 × 256 pixels and at most 16 megapixels. Refer to each image by position. “Image 1” is the first entry in images, “image 2” the second, and so on. Use the same words every time you mention an image. Specify what each image supplies. A reference can supply more than one feature, such as both the material and the background. Name those roles explicitly: Say how the pieces relate. Name positions and interactions, not only the parts: “the woman in image 2 sits on the swing in image 1, with the cat from image 3 on her lap.” Put the base image first. With aspect_ratio: "auto" (the default), the output takes its aspect ratio from image 1. Put the setting or base photo you are editing first, or set aspect_ratio explicitly when no reference has the shape you want.

Specify an object’s position

A written instruction such as “put the lamp from image 2 on the sideboard” lets FLUX 3 choose the position. To specify it, add box rows to the end of the prompt. A row with from: "ref_image_1" takes an element from the second image: src_bbox marks the object in image 2, and tgt_bbox marks where it belongs in the output. Anchor rows (from: "ref_image_0" with the same source and target box) ask the model to keep those elements in place.
The coordinates are illustrative: measure the boxes on your own images, as [top, left, bottom, right] on a 0–1000 grid. generate() is the helper from FLUX 3 Text to Image. Placement is a strong hint, not a hard constraint, and multi-reference compositions can place an element partly outside its box. Edit with boxes covers every row type.

More examples

These examples were generated with FLUX 3 Image. Hover an input to highlight where the prompt refers to it.

Paintings on coins, from four references

One image sets the scene and three supply artwork. Each painting appears on its own coin, and the hand and page stay the same.

A tower sinking into the sea

The tower is tilted and cut off at the waterline, with waves where it meets the sea.

A knife made of candy

The shape comes from image 2 and the material and background come from image 1, including the loose candies around the edge.

A heart in paper

Only the material changes. The shape, the veins, and the black background stay the same.

Use cases

The categories below cover the most common multi-reference tasks, each with the prompt that produced the result. The “Animal placed in scene”, “Pattern onto plate”, “Fill bottles with liquid”, and “Logo engraved in tree” results were generated with FLUX 3 Image. The rest were generated with FLUX.2.

Scene Compositing

Combine elements from multiple source images into a single coherent scene.
FLUX.2 result:
FLUX.2 result:

Style & Material Transfer

Apply the visual style, texture, or material of one image onto the content of another.
FLUX.2 result:
FLUX.2 result:

Object Replacement

Replace or fill objects with elements from another reference image.

Logo & Branding

Place logos from one image onto objects or scenes in another.
FLUX.2 result:

Checklist: vague to specific

Single-reference editing

Write the instruction for one image, with a catalog of edit types.

Edit with boxes

Place and anchor elements with box rows.