
GPT Image 2.5 Edit Prompts: Change One Thing, Keep the Rest
What GPT Image 2.5 edits really change, measured: a preserve list barely mattered, four chained edits nearly doubled the drift in her face, batching them halved it.
I have been writing thirty-clause preserve lists for two years: name the face, the pose, the light, or the model will change them. Last week I finally ran the same edit without one, and I could not tell the two results apart.
So I measured them. One photo, one change, two prompts: a five-word instruction, and an eighty-two-word version that spells out her face, her freckles, her hands, the mug, the notebook, the street, the camera angle and the colour temperature.
Inside a fixed box on her face, 0.3% of the pixels differed by more than 8 levels out of 255. The long list had bought me almost nothing.

That is the starting photo, made with the same model so there is nothing borrowed in this post. Everything below is an edit of it, or an edit of an edit.
A candid documentary photograph of a woman in her early thirties sitting at a
small outdoor cafe table on a quiet European city street. She wears a bright
red waxed-cotton jacket over a grey knit sweater. A plain white ceramic mug
sits on the table in front of her, next to a closed notebook. Late afternoon
natural light from the left, 35mm film look, eye-level camera, medium shot,
shallow depth of field, real skin texture with visible pores and a few
freckles, natural colour. No text, no logos, no watermarks.The GPT Image 2.5 edit prompt I thought I needed
Here is the careful version, the one with the preserve list:
Change one thing only: the red jacket becomes navy blue, same waxed-cotton
material and same cut. Keep everything else exactly as it is: her face,
facial features, skin tone, freckles, expression, hair, hands and pose; the
grey sweater; the white mug and the notebook; the table and the street behind
her; the camera angle, framing, depth of field, lighting direction and colour
temperature. Do not restyle the photograph, do not add or remove objects, and
do not add text, logos or watermarks.
It did the job. Her face is hers, the freckles are in the same places, the mug has not moved, the brown corduroy collar stayed brown, and the light still falls from the left.
The other version was five words: Make the jacket navy blue.
It produced the same picture, near enough.
What the preserve list actually changed: almost nothing
I measured it rather than eyeballing it. For each pair I took the mean absolute RGB difference per pixel, and the share of pixels that moved by more than 8 out of 255, inside a 240 by 210 pixel box on her face, at the same coordinates in every render.
The face is the part worth measuring because it is where a viewer notices drift first. It also means these numbers say nothing about the background. And I did not align the renders before comparing, so a render that shifts by a pixel or two pays for that shift in this number.
| Comparison | Mean change in the face box | Face pixels moved by more than 8/255 |
|---|---|---|
| Long preserve list vs five-word instruction | 1.37 | 0.3% |
| Five-word instruction vs the original photo | 4.75 | 13.4% |
| Long preserve list vs the original photo | 4.52 | 11.9% |
The two prompts differ from each other by less than either differs from the original. Against the original, the long list held 1.5 percentage points more of her face still than the five-word version did, 11.9% against 13.4%. That gap is not what separates a usable edit from a ruined one.
This is one scene, one subject and one model, so read it as a direction and not a benchmark. But it matches what I now see in day-to-day work: on this generation, keeping everything you did not mention is the default.
Every GPT Image 2.5 edit is a re-render, not a patch
Look at the second row of that table again. The five-word edit left 13.4% of the pixels in her face measurably different from the original, and I never mentioned her face.
Nothing is being pasted. The model reads your image, builds a new one that agrees with it, and hands that back. When the result looks untouched, it means the new render agrees with the old one closely enough that your eye cannot separate them. It does not mean the pixels survived.
You cannot treat an edit as a lossless local change, and the small disagreements accumulate every time you go round again.
Where the preserve list still earns its place
It is not dead. It earns its place in three situations. One of them showed up in these tests; the other two are why I still write the clause.
When the thing you named has fuzzy edges. Later in this post I asked for four changes in one prompt, jacket included. That run turned the brown corduroy collar navy as well, where both single-change runs had left it brown. The collar is part of the jacket when the model is repainting four things at once, and not part of it when it is repainting one. Naming the exception is the fix: the jacket becomes navy, keep the brown corduroy collar brown.
When something in the frame is easy to lose. Small props, reflections, a specific piece of text. Anything you would notice missing is worth one clause.
When the instruction could be read two ways. "Make it warmer" is a colour temperature to one reader and a wool coat to another.
The shape that works is short: what changes, then the exceptions that must not, then one catch-all line. Not the thirty-clause inventory I used to write.
Does GPT Image 2.5 have a mask or inpainting tool?
There is no mask tool in this workflow, and no brush. You do not paint over an area and then describe it, the way inpainting in a photo editor works. The only thing scoping the edit is the noun phrase you use, which is why the phrasing carries the weight.
So the skill worth practising is pointing at things in words. Two things usually do it: position, as in "the mug on the right of the table" rather than "the mug", or a property only that object has, as in "the dark red notebook". A relation works when neither is available: "the chair she is sitting on".
When two objects in the frame could match your words, assume the model will sometimes pick the wrong one. A more specific noun beats a longer sentence.
Removing an object with a GPT Image 2.5 edit prompt
Remove the white ceramic mug from the table. Everything else in the photograph
stays exactly as it is: the woman, her face, hair, jacket, pose and hands, the
notebook, the table surface and its reflections, the chair, the street behind
her, the camera angle, framing, lighting and colour. Fill the space where the
mug was with the table surface that belongs there. Do not add text, logos or
watermarks.
The mug is gone, the stone under it looks like stone, and the notebook and the red jacket came back unchanged to the eye, which per the section above is not the same as unchanged. The clause I would keep is the one about filling the space, because a removal is really a small generation: something has to be invented where the object was, and saying what belongs there is cheaper than fixing a smear afterwards. I did not run it without that sentence, so treat that as a habit rather than a measured effect.
Combining two images: give each reference a job
Up to 16 reference images go into one request. If you do not say what each one is for, you are gambling on which one dominates: I have had a product shot and a street scene come back blended into neither. Give each image a job in the first two lines.
Image 1 is the scene. Image 2 is the product. Place the amber bottle from
image 2 onto the cafe table in image 1, standing upright to the right of the
mug, at a size that matches the mug. It must sit in image 1's world: same late
afternoon light from the left, a matching soft shadow on the marble table, a
subtle reflection in the polished surface, same colour temperature and same
grain. Keep everything from image 1 unchanged.

The light is right: the highlight runs down the bottle's left side like everything else in the frame. The scale is not what I asked for. "A size that matches the mug" came back about 40% taller than the mug and set further back, which is a plausible bottle on a plausible table, but not the instruction.
The contact shadow is the weaker part. The mug darkens the stone to its right by about a third; beside the bottle the table reads the same brightness as the open surface beyond it. There is a dark line directly under the base and nothing cast to the side. Contact shadows are the first thing to check on any composite, because a floating object is the giveaway that survives everything else looking correct.
Locking text through a GPT Image 2.5 edit
Text is where edits used to fall apart. The rule that works is to quote the line, say it appears once, and describe the type.
Put the bottle from the input image on a large roadside billboard photographed
at sunset, seen from the road at a slight angle. The billboard carries exactly
one line of copy: "SLOW MORNINGS". Set it in a bold sans-serif, evenly spaced,
high contrast against the billboard background, easy to read, and appearing
once only. The bottle keeps its shape, colour and matte black cap. No other
text, no logos, no watermarks, no brand marks on the billboard, the frame or
the surroundings.
Two letters wrong is a reshoot, so this is worth reading character by character before you use it. Here the line is correct, appears once, and nothing else crept onto the board. Note also that the output is 3:2 while the input bottle shot was 2:3: an edit takes the aspect ratio you ask for, and reframes to fit it.
Then the second round, which is one sentence:
Make it a winter evening with snow falling. Keep the billboard, the bottle,
the line of copy and its typography, the camera angle and the composition
exactly as they are.
Season changed, sign did not. Snow settled on the frame and the lamps read brighter against the darker sky, although they were already lit at sunset, and the letters kept their spacing. That is chaining working exactly as it should, and it is the case for editing in rounds: I could not have described that snow scene from scratch as precisely as I could point at the sunset one and change the weather. It also helps that there is no face in that frame. The next section is what rounds cost when there is one.
Chaining edits: what four rounds cost me
Chaining has a price, and I wanted a number for it. Starting from the original photo I ran four edits in sequence, each a single short instruction, none of them about her face:
Make the jacket navy blue.Change the sweater to cream.Change the notebook on the table to a dark red one.Change the mug to a dark green one.
Round one is the five-word run from the first table, reused as the chain's starting point, which is why its numbers match.
| Rounds of editing | Mean change in the face box | Face pixels moved by more than 8/255 |
|---|---|---|
| 1 | 4.75 | 13.4% |
| 2 | 9.02 | 42.8% |
| 3 | 11.46 | 53.1% |
| 4 | 12.45 | 56.1% |

Every change I asked for landed. She is also not quite the same photograph any more. The freckles are heavier, the street behind her has rearranged a few details, and at full size the skin has picked up a faint patterned texture that is not freckles. Still the same woman at a glance, and no longer a picture I could put next to the original in a set.
The jump from round one to round two is the part worth remembering: the second pass more than triples the number of face pixels that move. The curve is steepest at the start, and nothing in the output tells you where on it you are.
Batch four changes into one prompt instead of four rounds
So I went back to the original photo and asked for all four changes in a single prompt instead.
| Same four changes | Mean change in the face box | Face pixels moved by more than 8/255 |
|---|---|---|
| Four rounds, one change each | 12.45 | 56.1% |
| One round, four changes | 6.86 | 31.2% |

Half the drift for the same result. This cuts against the advice everyone repeats, including advice I have given: one change per turn is a good habit when you are still deciding, and an expensive one when you already know what you want. Decide first, then write one prompt.
The trade is precision at the edges, and it is the collar case from earlier: the batched run swept the corduroy collar into the jacket change, where the single-change run had left it brown. Batch the changes, and name the exceptions you care about.
When you do need several rounds, restart from the original as soon as you can, and carry over what you learned in the wording rather than in the image.
Editing a photo you uploaded yourself
Nothing above depends on the image having been generated. The same prompt shape works on a photo you upload to the image-to-image generator: describe the change, name the exceptions, add the catch-all line.
One consequence of the re-render finding is worth saying plainly before you edit something irreplaceable. The file you get back is a new image, not your original with a patch applied. Every pixel has been through the model, including the ones in the corner you never mentioned. For a product shot or a mockup that is fine. For the only copy of a photograph of somebody, keep the original and treat the edit as a derived file.
The practical limits are the same as for any other render: the output arrives at the ratio and resolution you asked for, not at your file's dimensions, so a 4000 pixel wide photo edited at 1K comes back at 1K. Crop and resize planning belongs after the edit, not before it.
What four rounds did not move: the framing
I expected the camera to creep, and went looking for it. It is not there. Measuring the dark span of her hair across three rows, her head is 457, 458 and 498 pixels wide in the original and 455, 456 and 496 by round four: the same size, inside the noise. The whole frame has shifted right by two or three pixels and nothing has scaled.
So what four rounds cost is texture and colour, not composition. That is worth knowing if you are building a set, because it tells you which mismatch you will be looking at: a chained render will line up with its source and still not match it. You cannot fix that by cropping.
It also means the numbers above are slightly pessimistic. A fixed box compared without alignment charges those two or three pixels of shift to the drift column, so the real texture change is a little smaller than the table says.
Translating labels without moving the layout
A useful edit that has nothing to do with photography: take a finished diagram and swap the language.

Translate every piece of text in this diagram into Spanish. Nothing else
changes: same layout, same icons, same arrows, same colours, same type sizes
and same positions. Only the words are replaced.
Serpentín exterior, Compresor, Válvula de expansión, Serpentín interior, and the title above them. Accents correct, icons identical, arrows in the same places, the blue and orange loop untouched.
One thing did move. "Warm air to rooms" was two lines at one size; the Spanish "Aire cálido a las habitaciones" is longer, and the model set the second line smaller to fit. Longer languages push type around, so check line breaks and type sizes in the translated version rather than assuming a clean swap. German compounds run longer again, so I would expect it to push harder, but I did not test it.
Settings for GPT Image 2.5 edits
| Setting | What I used | What it did |
|---|---|---|
| Model | Flare | Fast enough to run a chain of four in a few minutes. Sunburst is the one to switch to for fine detail work, and the Flare vs Sunburst tests show where it pays off |
| Aspect ratio | 2:3 for the cafe photo, 3:2 for the billboard and the diagram | The output takes the ratio you request, not the input's, and reframes to fit |
| Resolution | 1K | Renders came back at 832×1248 and 1248×832 |
| Quality | medium for the photo edits, high where text had to be exact | Every text render in this post was at high |
| Reference images | 1 for edits, 2 for the composite | Up to 16 per request, each one given a job |
| Credits | 8 per render at medium and 1K, 32 at high and 1K | The photo edits cost 8 each, the billboard and diagram edits 32. The four-round chain cost four renders, the batched version one |
Four ways a GPT Image 2.5 edit goes wrong
| What you see | Why it happens | The fix |
|---|---|---|
| Something next to your target changed too | Your noun covers more than you meant. "The jacket" includes the collar | Name the exception in the same sentence |
| The person looks subtly different after a few rounds | Each round re-renders the whole frame and small disagreements stack | Go back to the original and batch, or accept the drift and stop calling it the same photo |
| A composited object floats | The contact shadow is too light, even when scale and colour are right | Asking for the shadow is not enough: mine asked and still came back light. Check where the object meets the surface, and spend a round on that alone |
| Translated text reflows | The new language is longer and type sizes adapt to fit | Check line breaks and sizes, not just spelling |
What I did not test
The numbers in this post come from one scene, one subject, Flare, medium quality and 1K. Four things stayed outside it, and I would not assume the results carry over:
- Sunburst. The precision model may hold a face tighter across a chain. I measured Flare only.
- 2K and high quality on photo edits. Every photo edit here ran at medium and 1K. More pixels may mean more or less drift; I have no data either way.
- Faces I did not generate. A real person's photograph is the case where drift matters most, and it is also the case where I would not publish a test.
- Long chains. I stopped at four rounds. The curve was still rising, and the interesting question is where it flattens.
A three-step routine for GPT Image 2.5 edits
- Write down every change you want before you touch the model. If the list is settled, it belongs in one prompt.
- Write the prompt as: what changes, the exceptions that stay, one catch-all line. Quote any text that must appear, and say it appears once.
- Compare the result against the original at full size, in this order: the thing you changed, the edges around it, faces and text, then everything else. If you need another round, decide whether to continue from this render or restart from the original.
You can run all of this in the browser from the image-to-image generator, or from the GPT Image 2.5 page if you want the model controls in front of you.
The Bottom Line
The preserve list is no longer what saves an edit on this model: a five-word prompt and an eighty-two-word one produced faces that differed by 0.3% of their pixels. What costs you is rounds. Four single-change edits moved her face nearly twice as far as the same four changes asked for at once, and nothing in those prompts ever mentioned her face. Decide the whole change before you start, put it in one prompt, name the exceptions you care about, and go back to the original rather than editing the edit.
Related reading
- GPT Image 2.5 Flare vs Sunburst: 6 Same-Prompt Tests
- GPT Image 2 Prompt Writing Guide: 7 Rules for 90% Hit Rate
- GPT Image 2.5 Product Photo Prompts: 10 E-commerce Shots
- GPT Image 2.5 for Scientific Figures: What Journals Accept
gpt-image2.art is an independent site and is not affiliated with OpenAI. Every render in this post was generated and edited with GPT Image 2.5 Flare on 24 September 2026; the measurements come from one scene at 1K, so treat them as a direction rather than a benchmark.
他の記事

GPT Image 2.5 Product Photo Prompts: 10 E-commerce Shots
GPT Image 2.5 product photo prompts for online sellers: white background, lifestyle, flat lay, infographic, banner and more, each run on Flare and Sunburst.

GPT Image 2.5 for Scientific Figures: What Journals Accept
GPT Image 2.5 scientific figures: which types journals accept, the style preamble to reuse on every flow diagram, pipeline and architecture figure, and the errors reviewers catch.

10 Cinematic Camera Shot Prompts for GPT Image 2
10 copy-ready cinematic camera shot prompts for GPT Image 2, plus a practical shot formula, continuity workflow, and composition fixes.
Generate your first image with GPT Image 2 — right now
Reliable non-Latin text rendering, directed editing, and 50+ ready-to-use prompts. No downloads — just open in your browser.