GPT Image 2.5 Edit Prompts: Change One Thing, Keep the Rest
2026/09/25

GPT Image 2.5 Edit Prompts: Change One Thing, Keep the Rest

What GPT Image 2.5 edits really change, measured: a preserve list barely mattered, four chained edits nearly doubled the drift in her face, batching them halved it.

I have been writing thirty-clause preserve lists for two years: name the face, the pose, the light, or the model will change them. Last week I finally ran the same edit without one, and I could not tell the two results apart.

So I measured them. One photo, one change, two prompts: a five-word instruction, and an eighty-two-word version that spells out her face, her freckles, her hands, the mug, the notebook, the street, the camera angle and the colour temperature.

Inside a fixed box on her face, 0.3% of the pixels differed by more than 8 levels out of 255. The long list had bought me almost nothing.

A woman in a red waxed-cotton jacket at an outdoor cafe table on a cobbled street, a white mug and a closed notebook in front of her

That is the starting photo, made with the same model so there is nothing borrowed in this post. Everything below is an edit of it, or an edit of an edit.

A candid documentary photograph of a woman in her early thirties sitting at a
small outdoor cafe table on a quiet European city street. She wears a bright
red waxed-cotton jacket over a grey knit sweater. A plain white ceramic mug
sits on the table in front of her, next to a closed notebook. Late afternoon
natural light from the left, 35mm film look, eye-level camera, medium shot,
shallow depth of field, real skin texture with visible pores and a few
freckles, natural colour. No text, no logos, no watermarks.

The GPT Image 2.5 edit prompt I thought I needed

Here is the careful version, the one with the preserve list:

Change one thing only: the red jacket becomes navy blue, same waxed-cotton
material and same cut. Keep everything else exactly as it is: her face,
facial features, skin tone, freckles, expression, hair, hands and pose; the
grey sweater; the white mug and the notebook; the table and the street behind
her; the camera angle, framing, depth of field, lighting direction and colour
temperature. Do not restyle the photograph, do not add or remove objects, and
do not add text, logos or watermarks.

The same photo and the same woman, with the jacket now navy blue and the brown corduroy collar unchanged

It did the job. Her face is hers, the freckles are in the same places, the mug has not moved, the brown corduroy collar stayed brown, and the light still falls from the left.

The other version was five words: Make the jacket navy blue.

It produced the same picture, near enough.

What the preserve list actually changed: almost nothing

I measured it rather than eyeballing it. For each pair I took the mean absolute RGB difference per pixel, and the share of pixels that moved by more than 8 out of 255, inside a 240 by 210 pixel box on her face, at the same coordinates in every render.

The face is the part worth measuring because it is where a viewer notices drift first. It also means these numbers say nothing about the background. And I did not align the renders before comparing, so a render that shifts by a pixel or two pays for that shift in this number.

ComparisonMean change in the face boxFace pixels moved by more than 8/255
Long preserve list vs five-word instruction1.370.3%
Five-word instruction vs the original photo4.7513.4%
Long preserve list vs the original photo4.5211.9%

The two prompts differ from each other by less than either differs from the original. Against the original, the long list held 1.5 percentage points more of her face still than the five-word version did, 11.9% against 13.4%. That gap is not what separates a usable edit from a ruined one.

This is one scene, one subject and one model, so read it as a direction and not a benchmark. But it matches what I now see in day-to-day work: on this generation, keeping everything you did not mention is the default.

Every GPT Image 2.5 edit is a re-render, not a patch

Look at the second row of that table again. The five-word edit left 13.4% of the pixels in her face measurably different from the original, and I never mentioned her face.

Nothing is being pasted. The model reads your image, builds a new one that agrees with it, and hands that back. When the result looks untouched, it means the new render agrees with the old one closely enough that your eye cannot separate them. It does not mean the pixels survived.

You cannot treat an edit as a lossless local change, and the small disagreements accumulate every time you go round again.

Where the preserve list still earns its place

It is not dead. It earns its place in three situations. One of them showed up in these tests; the other two are why I still write the clause.

When the thing you named has fuzzy edges. Later in this post I asked for four changes in one prompt, jacket included. That run turned the brown corduroy collar navy as well, where both single-change runs had left it brown. The collar is part of the jacket when the model is repainting four things at once, and not part of it when it is repainting one. Naming the exception is the fix: the jacket becomes navy, keep the brown corduroy collar brown.

When something in the frame is easy to lose. Small props, reflections, a specific piece of text. Anything you would notice missing is worth one clause.

When the instruction could be read two ways. "Make it warmer" is a colour temperature to one reader and a wool coat to another.

The shape that works is short: what changes, then the exceptions that must not, then one catch-all line. Not the thirty-clause inventory I used to write.

Does GPT Image 2.5 have a mask or inpainting tool?

There is no mask tool in this workflow, and no brush. You do not paint over an area and then describe it, the way inpainting in a photo editor works. The only thing scoping the edit is the noun phrase you use, which is why the phrasing carries the weight.

So the skill worth practising is pointing at things in words. Two things usually do it: position, as in "the mug on the right of the table" rather than "the mug", or a property only that object has, as in "the dark red notebook". A relation works when neither is available: "the chair she is sitting on".

When two objects in the frame could match your words, assume the model will sometimes pick the wrong one. A more specific noun beats a longer sentence.

Removing an object with a GPT Image 2.5 edit prompt

Remove the white ceramic mug from the table. Everything else in the photograph
stays exactly as it is: the woman, her face, hair, jacket, pose and hands, the
notebook, the table surface and its reflections, the chair, the street behind
her, the camera angle, framing, lighting and colour. Fill the space where the
mug was with the table surface that belongs there. Do not add text, logos or
watermarks.

The same cafe photo with the white mug gone and the marble table surface filled in where it stood

The mug is gone, the stone under it looks like stone, and the notebook and the red jacket came back unchanged to the eye, which per the section above is not the same as unchanged. The clause I would keep is the one about filling the space, because a removal is really a small generation: something has to be invented where the object was, and saying what belongs there is cheaper than fixing a smear afterwards. I did not run it without that sentence, so treat that as a habit rather than a measured effect.

Combining two images: give each reference a job

Up to 16 reference images go into one request. If you do not say what each one is for, you are gambling on which one dominates: I have had a product shot and a street scene come back blended into neither. Give each image a job in the first two lines.

Image 1 is the scene. Image 2 is the product. Place the amber bottle from
image 2 onto the cafe table in image 1, standing upright to the right of the
mug, at a size that matches the mug. It must sit in image 1's world: same late
afternoon light from the left, a matching soft shadow on the marble table, a
subtle reflection in the polished surface, same colour temperature and same
grain. Keep everything from image 1 unchanged.

An unbranded amber glass bottle with a matte black pump cap on a light grey studio background

The cafe photo with an amber pump bottle standing on the table to the right of the mug, in front of the notebook

The light is right: the highlight runs down the bottle's left side like everything else in the frame. The scale is not what I asked for. "A size that matches the mug" came back about 40% taller than the mug and set further back, which is a plausible bottle on a plausible table, but not the instruction.

The contact shadow is the weaker part. The mug darkens the stone to its right by about a third; beside the bottle the table reads the same brightness as the open surface beyond it. There is a dark line directly under the base and nothing cast to the side. Contact shadows are the first thing to check on any composite, because a floating object is the giveaway that survives everything else looking correct.

Locking text through a GPT Image 2.5 edit

Text is where edits used to fall apart. The rule that works is to quote the line, say it appears once, and describe the type.

Put the bottle from the input image on a large roadside billboard photographed
at sunset, seen from the road at a slight angle. The billboard carries exactly
one line of copy: "SLOW MORNINGS". Set it in a bold sans-serif, evenly spaced,
high contrast against the billboard background, easy to read, and appearing
once only. The bottle keeps its shape, colour and matte black cap. No other
text, no logos, no watermarks, no brand marks on the billboard, the frame or
the surroundings.

A roadside billboard at sunset showing the amber bottle and the line SLOW MORNINGS in bold sans-serif type

Two letters wrong is a reshoot, so this is worth reading character by character before you use it. Here the line is correct, appears once, and nothing else crept onto the board. Note also that the output is 3:2 while the input bottle shot was 2:3: an edit takes the aspect ratio you ask for, and reframes to fit it.

Then the second round, which is one sentence:

Make it a winter evening with snow falling. Keep the billboard, the bottle,
the line of copy and its typography, the camera angle and the composition
exactly as they are.

The same billboard at dusk in falling snow, the sun still low on the horizon, with the same bottle and the same line of copy intact

Season changed, sign did not. Snow settled on the frame and the lamps read brighter against the darker sky, although they were already lit at sunset, and the letters kept their spacing. That is chaining working exactly as it should, and it is the case for editing in rounds: I could not have described that snow scene from scratch as precisely as I could point at the sunset one and change the weather. It also helps that there is no face in that frame. The next section is what rounds cost when there is one.

Chaining edits: what four rounds cost me

Chaining has a price, and I wanted a number for it. Starting from the original photo I ran four edits in sequence, each a single short instruction, none of them about her face:

  1. Make the jacket navy blue.
  2. Change the sweater to cream.
  3. Change the notebook on the table to a dark red one.
  4. Change the mug to a dark green one.

Round one is the five-word run from the first table, reused as the chain's starting point, which is why its numbers match.

Rounds of editingMean change in the face boxFace pixels moved by more than 8/255
14.7513.4%
29.0242.8%
311.4653.1%
412.4556.1%

The cafe photo after four chained edits: navy jacket, cream sweater, dark red notebook, dark green mug

Every change I asked for landed. She is also not quite the same photograph any more. The freckles are heavier, the street behind her has rearranged a few details, and at full size the skin has picked up a faint patterned texture that is not freckles. Still the same woman at a glance, and no longer a picture I could put next to the original in a set.

The jump from round one to round two is the part worth remembering: the second pass more than triples the number of face pixels that move. The curve is steepest at the start, and nothing in the output tells you where on it you are.

Batch four changes into one prompt instead of four rounds

So I went back to the original photo and asked for all four changes in a single prompt instead.

Same four changesMean change in the face boxFace pixels moved by more than 8/255
Four rounds, one change each12.4556.1%
One round, four changes6.8631.2%

The cafe photo with all four changes made in a single edit: navy jacket, cream sweater, dark red notebook, dark green mug

Half the drift for the same result. This cuts against the advice everyone repeats, including advice I have given: one change per turn is a good habit when you are still deciding, and an expensive one when you already know what you want. Decide first, then write one prompt.

The trade is precision at the edges, and it is the collar case from earlier: the batched run swept the corduroy collar into the jacket change, where the single-change run had left it brown. Batch the changes, and name the exceptions you care about.

When you do need several rounds, restart from the original as soon as you can, and carry over what you learned in the wording rather than in the image.

Editing a photo you uploaded yourself

Nothing above depends on the image having been generated. The same prompt shape works on a photo you upload to the image-to-image generator: describe the change, name the exceptions, add the catch-all line.

One consequence of the re-render finding is worth saying plainly before you edit something irreplaceable. The file you get back is a new image, not your original with a patch applied. Every pixel has been through the model, including the ones in the corner you never mentioned. For a product shot or a mockup that is fine. For the only copy of a photograph of somebody, keep the original and treat the edit as a derived file.

The practical limits are the same as for any other render: the output arrives at the ratio and resolution you asked for, not at your file's dimensions, so a 4000 pixel wide photo edited at 1K comes back at 1K. Crop and resize planning belongs after the edit, not before it.

What four rounds did not move: the framing

I expected the camera to creep, and went looking for it. It is not there. Measuring the dark span of her hair across three rows, her head is 457, 458 and 498 pixels wide in the original and 455, 456 and 496 by round four: the same size, inside the noise. The whole frame has shifted right by two or three pixels and nothing has scaled.

So what four rounds cost is texture and colour, not composition. That is worth knowing if you are building a set, because it tells you which mismatch you will be looking at: a chained render will line up with its source and still not match it. You cannot fix that by cropping.

It also means the numbers above are slightly pessimistic. A fixed box compared without alignment charges those two or three pixels of shift to the drift column, so the real texture change is a little smaller than the table says.

Translating labels without moving the layout

A useful edit that has nothing to do with photography: take a finished diagram and swap the language.

A flat vector diagram titled How a heat pump heats a house, with five English labels

Translate every piece of text in this diagram into Spanish. Nothing else
changes: same layout, same icons, same arrows, same colours, same type sizes
and same positions. Only the words are replaced.

The same diagram with every label in Spanish, identical icons, arrows and colours

Serpentín exterior, Compresor, Válvula de expansión, Serpentín interior, and the title above them. Accents correct, icons identical, arrows in the same places, the blue and orange loop untouched.

One thing did move. "Warm air to rooms" was two lines at one size; the Spanish "Aire cálido a las habitaciones" is longer, and the model set the second line smaller to fit. Longer languages push type around, so check line breaks and type sizes in the translated version rather than assuming a clean swap. German compounds run longer again, so I would expect it to push harder, but I did not test it.

Settings for GPT Image 2.5 edits

SettingWhat I usedWhat it did
ModelFlareFast enough to run a chain of four in a few minutes. Sunburst is the one to switch to for fine detail work, and the Flare vs Sunburst tests show where it pays off
Aspect ratio2:3 for the cafe photo, 3:2 for the billboard and the diagramThe output takes the ratio you request, not the input's, and reframes to fit
Resolution1KRenders came back at 832×1248 and 1248×832
Qualitymedium for the photo edits, high where text had to be exactEvery text render in this post was at high
Reference images1 for edits, 2 for the compositeUp to 16 per request, each one given a job
Credits8 per render at medium and 1K, 32 at high and 1KThe photo edits cost 8 each, the billboard and diagram edits 32. The four-round chain cost four renders, the batched version one

Four ways a GPT Image 2.5 edit goes wrong

What you seeWhy it happensThe fix
Something next to your target changed tooYour noun covers more than you meant. "The jacket" includes the collarName the exception in the same sentence
The person looks subtly different after a few roundsEach round re-renders the whole frame and small disagreements stackGo back to the original and batch, or accept the drift and stop calling it the same photo
A composited object floatsThe contact shadow is too light, even when scale and colour are rightAsking for the shadow is not enough: mine asked and still came back light. Check where the object meets the surface, and spend a round on that alone
Translated text reflowsThe new language is longer and type sizes adapt to fitCheck line breaks and sizes, not just spelling

What I did not test

The numbers in this post come from one scene, one subject, Flare, medium quality and 1K. Four things stayed outside it, and I would not assume the results carry over:

  • Sunburst. The precision model may hold a face tighter across a chain. I measured Flare only.
  • 2K and high quality on photo edits. Every photo edit here ran at medium and 1K. More pixels may mean more or less drift; I have no data either way.
  • Faces I did not generate. A real person's photograph is the case where drift matters most, and it is also the case where I would not publish a test.
  • Long chains. I stopped at four rounds. The curve was still rising, and the interesting question is where it flattens.

A three-step routine for GPT Image 2.5 edits

  1. Write down every change you want before you touch the model. If the list is settled, it belongs in one prompt.
  2. Write the prompt as: what changes, the exceptions that stay, one catch-all line. Quote any text that must appear, and say it appears once.
  3. Compare the result against the original at full size, in this order: the thing you changed, the edges around it, faces and text, then everything else. If you need another round, decide whether to continue from this render or restart from the original.

You can run all of this in the browser from the image-to-image generator, or from the GPT Image 2.5 page if you want the model controls in front of you.

The Bottom Line

The preserve list is no longer what saves an edit on this model: a five-word prompt and an eighty-two-word one produced faces that differed by 0.3% of their pixels. What costs you is rounds. Four single-change edits moved her face nearly twice as far as the same four changes asked for at once, and nothing in those prompts ever mentioned her face. Decide the whole change before you start, put it in one prompt, name the exceptions you care about, and go back to the original rather than editing the edit.

gpt-image2.art is an independent site and is not affiliated with OpenAI. Every render in this post was generated and edited with GPT Image 2.5 Flare on 24 September 2026; the measurements come from one scene at 1K, so treat them as a direction rather than a benchmark.

Free to try

Generate your first image with GPT Image 2 — right now

Reliable non-Latin text rendering, directed editing, and 50+ ready-to-use prompts. No downloads — just open in your browser.