Skip to main content
A layout prompt combines a caption with a JSON array of elements. The caption describes the scene and refers to each element by <id>. Each array entry supplies its position in bbox and appearance in desc. For the coordinate format and requests, see FLUX 3 Image Editing.

Write the caption

Start with the medium, setting, and light. Introduce the subjects and describe how they relate, placing each <id> next to the element it names:
Describe positions in words as well as coordinates. “In the left foreground” and “in her lap” help explain the arrangement represented by the boxes.

Describe each element

Use desc for appearance: material, color, pose, expression, and visible details. Keep names short and unique, such as person_1 and person_2.
Give each text block a row. Quote the words and specify case, color, type style, and alignment. Use \n for line breaks:

Size the boxes

Choose aspect_ratio before placing boxes. Match each box to the intended extent of its element. Overlap boxes when objects overlap; put a dog’s box inside its owner’s when it sits in their lap. Use one box for a distant crowd unless individual positions matter. A background can cover the full frame, [0, 0, 1000, 1000], or only the area where it appears, such as the sky above a crowd.
Grid units are relative to each axis. On a 2:3 image, a square needs a box 150 units wide for every 100 units tall.

Explore three layouts

Select a row or caption token to see the corresponding box. These are the original prompts and results from the FLUX 3 Image launch examples.
  • Le Festival du Soleil: the title has its own box in the sky above the dome. The crowd and swimmers each use a collection box.
  • Two tall panels: the panel boxes contain smaller boxes for the figures, landscape, and text. The caption describes how they fit together.
  • Winter spectators: each dog’s box overlaps its owner’s. The background descriptions specify blur, while the foreground descriptions include detail.

Start from a short prompt

Each example below pairs a one-sentence idea with the caption and boxes used to generate the image. Start from one of these layouts and adapt its descriptions and coordinates to your scene. Compare each sentence with its layout to see what a full layout adds:
  • Dinner by the garden: “two diners” becomes one box per diner. The layout adds the table, dishes, and a candle inside, and the pine tree, fence, path, and lights in the garden.
  • Tools, the magazine: the sentence names the title and the plank tower. The layout gives each its own box and adds a subtitle, a price, and a barcode, as on a real cover.
  • Sixteen layouts: one sentence becomes 32 boxes, two for each poster in the 4×4 grid.
To change a planned layout, edit its boxes or descriptions and send it again.

Adjust the result

If an element appears in the wrong place, check that its caption token matches its id, its coordinates use y before x, and its size suits the frame. Describe the intended position in the caption too. Boxes guide composition rather than clipping content. To change an existing image, use source and target boxes and state what should change and what should stay.