
GPT Image 2.5 3D Map Prompts: Cities as Miniature Dioramas
A GPT Image 2.5 3D map prompt that grows a miniature city out of a real map. One template, four cities, and the accuracy checks that decide whether an AI 3D city map is usable.
The first one I made was Kyoto, and I nearly missed the interesting part. I was looking at the pagoda and the temple roofs, and only later noticed that the map underneath said Lake Biwa in the right place, with 琵琶湖 set under it in smaller type, and that Osaka Bay was where Osaka Bay goes.
The miniature city is the part people react to, but the map is the part that makes it look real. Get it wrong and you have a cute model floating on a green smear. Get it right and it reads like a photograph of something in a museum case.
This is one reusable GPT Image 2.5 3D map prompt with six slots in it: a city, a region, and four landmark lines. Below it are the checks I run before I use any of the output. I built four cities with it, and it was startlingly accurate in places and confidently wrong in others.

The GPT Image 2.5 3D map prompt template
Everything below is one prompt with six slots. Fill in [CITY] and [REGION], then the four landmark lines from the table in the next section, and leave the rest alone the first time.
A square, highly detailed image.
The base is a realistic topographic map seen from directly above,
accurately showing the terrain, coastline, rivers and roads of
[REGION], with clean natural cartographic labelling texture.
Rising from the true geographic position of [CITY] on that map is a
miniature three-dimensional cityscape, like an exquisite 3D diorama
model growing naturally out of the map surface.
The miniature city blends [CITY]'s most recognisable elements:
[LANDMARK 1], [LANDMARK 2], [TERRAIN FEATURE], [STREET-LEVEL DETAIL].
These elements connect naturally with the map beneath, appearing to
emerge from it rather than sit on top.
Realistic, refined, collector-grade model quality, ultra-real
miniature photography. Emphasise true materials, fine texture,
layered depth and miniature scale, with shallow depth of field so the
city is sharp and the surrounding map softly blurred. Soft studio
light with natural shadows, clean composition, cinematic and premium
travel visual.
No brand logos, no watermarks, no people.Four lines in there are load-bearing. Those are the ones not to delete once you start editing.
"Rising from the true geographic position" is what aims the city at its own label instead of the middle of the frame. With it in place, the city lands near where the label goes and the map bends around it.
"Growing naturally out of the map surface" and the follow-up sentence about emerging rather than sitting on top say the same thing twice on purpose. It is the instruction I would expect to get dropped first, and the failure it guards against is a model sitting on a map like a chess piece on a board, with a visible seam.
"Shallow depth of field so the city is sharp and the surrounding map softly blurred" is what sells the miniature illusion. Real photography of small objects has a very narrow plane of focus, and our eyes read that blur as this thing is tiny. Without it you are back to a video-game map look.
"No brand logos, no watermarks, no people" is not optional. Cities are dense with signage, and this is the genre where an invented storefront mark is most likely to appear. This line stays on every prompt here.
Fill the four landmark slots in the GPT Image 2.5 prompt
The line that decides whether a city is recognisable is the landmark list. You have four slots. Use them like this rather than writing "iconic buildings":
| Slot | What goes in it | Kyoto example |
|---|---|---|
| Landmark 1 | The single silhouette people would name | a tall pagoda |
| Landmark 2 | A second structure with a different shape | traditional timber temple roofs |
| Terrain feature | What the city sits in or on | forested hills |
| Street-level detail | Something small that sets the scale | a river through low machiya townhouses |
The street-level detail matters more than it looks. A pagoda alone gives you a postcard. A pagoda plus low two-storey houses gives the viewer something to measure the pagoda against, and that comparison is what sells the scale.
One note on naming: I described every landmark by its shape rather than naming it, and it worked. "A flat-topped mountain with a cable line" produced a clean Table Mountain. "A tall ridged concrete church tower" produced something very close to Hallgrímskirkja. I did not run a named-building comparison, but naming a building that is also a trademark is a problem you do not need. Describe the silhouette.
Four cities from one GPT Image 2.5 3D map prompt
All four ran on GPT Image 2.5 Flare at 1:1, first attempt, no retries, so this is the base rate rather than a best-of.
Lisbon, Portugal
Landmark 1: a long red suspension bridge over the water
Landmark 2: a hilltop castle
Terrain: terracotta-roofed hillside houses stacked on steep slopes
Detail: a yellow funicular tram climbing a narrow street
The labels are right, and in Portuguese: Rio Tejo, Oceano Atlântico with the circumflex on the â, Almada, Seixal, Setúbal with its accent, Cascais and Sintra out west. The motorway shields read A5, A2 and A33, which are real roads in roughly the right places. The tram is yellow, the bridge is red, the castle is on the hill.
The geography is wrong. Real Lisbon sits on the north bank of the Tagus, and the bridge crosses south. Here the city has been given a peninsula that does not exist and the estuary reshaped around it. The names are right, the relationships between the names are roughly right, and the coastline is invented.
Correct labels on a wrong shape turned out to be the thing to watch for.
Cape Town, South Africa
Landmark 1: a flat-topped mountain with a cable line
Landmark 2: a curving harbour waterfront
Terrain: coastal cliffs meeting the sea
Detail: low pastel-coloured terraced houses
This one surprised me. It labelled Table Mountain at 1086 m and Lion's Head at 669 m. Both are correct. It put the Atlantic on one side and the Indian Ocean on the other, wrote Western Cape across the interior, and ran an N2 shield along the road east.
So within one batch, on one model, one image invented a coastline and another gave two peak elevations accurate to the metre. You cannot predict which you will get, which is the entire argument for the checklist further down.
Reykjavík, Iceland
Landmark 1: a tall ridged concrete church tower
Landmark 2: a glass-faceted waterfront concert hall
Terrain: steaming geothermal vents at the edge of town
Detail: brightly painted corrugated-metal rooftops
The strongest of the four on text. Faxaflói, Seltjarnarnes, Kópavogur, Garðabær, Mosfellsbær, Laugardalur, Reykjanesfólkvangur: every one spelled correctly, including the eth, the ash and the accented vowels. Esja is marked 914 m, which is right. Route 1 shields mark the ring road.
If you have been avoiding accented or non-Latin place names because you expected mangled characters, this genre is worth retesting.
Kyoto, Japan
The first image in this post. Bilingual labels throughout: Lake Biwa with 琵琶湖, Kobe with 神戸, Osaka with 大阪, plus Nara, Shiga, Hyōgo and Wakayama as prefectures in both scripts, arranged roughly correctly relative to each other.
The macrons survived, which I did not expect. Hyōgo and Kōbe both carry a proper macron rather than the dieresis these usually degrade into.
The defect is one layer down. The small secondary labels scattered across the map, the ones at the size of a minor road name, are not words at all. They are kanji-shaped marks: right stroke density, right rhythm, not characters. At full size they are obviously wrong; scaled down to a social post they read as map texture and nobody notices.
That split held across all four cities: large labels correct down to the diacritics, small labels not language at all. It is the most useful thing to know before you publish one.
What the AI 3D city map gets right, and what it invents
Four images is not a benchmark, but the pattern across them was consistent enough to act on.
Reliable: major place names and their spelling, including accents and CJK characters. The relative arrangement of named places. Landmark silhouettes when described by shape. Road shield styling. The miniature-photography look itself, which it lands essentially every time.
Unreliable: coastlines and landmass shapes. The exact position of the city relative to its own label. Any text below roughly the size of a city name, which was not language at all in the one image where I looked closely enough to tell. Elevation figures. Four appear across the four maps. I checked three of them (1086 m, 669 m, 914 m) and all three were right. The fourth is Bláfjöll, marked 923 m in the blurred foreground of the Reykjavík map, and the sources I could find disagree with each other and with the figure. That is the point: a number you cannot quickly settle is a number you should not ship.
The useful way to hold this: the model is drawing something that looks like a map of that place, not reading a map of that place. It has absorbed what the Kansai region looks like as an image. It is not consulting geography. Where that memory is strong the labels are facts. Where it is weaker it fills the gap with something that looks right.
Three checks before you use an AI map diorama
Three passes, in this order, because each one can send you back to regenerate and there is no point doing the fine check first.
Pass one, shape. Compare the coastline and any major water feature against a real map. If the landmass is wrong in a way a local would notice, regenerate rather than edit. Shape errors are baked into the whole composition.
Pass two, text. Zoom to full size and read every label you can read. Check spelling, check diacritics, check that named places sit sensibly relative to each other. Then decide what to do about any gibberish small labels. Kyoto had them; on the other three I could not see any at a size that would tell. For social use, leave them. For anything printed or client-facing, crop them out or ask for a version with fewer labels.
Pass three, numbers and logos. Any elevation, distance or population on the map is a claim. Verify it or remove it. Then scan the miniature city for invented signage. This is the pass people skip. Nothing in my four runs had an invented logo in it, but the negative clause was on every time, and a dense cityscape is exactly where one would turn up.
If you want the map decorative rather than factual, there is a shortcut: add use plausible but unlabelled cartographic styling, no place names to the prompt. You lose the thing that makes these feel real, but you also lose every failure mode in passes two and three.
GPT Image 2.5 settings and specs for map dioramas
These are the settings I start from and what came back on the files:
| Setting | What I use | What I measured |
|---|---|---|
| Aspect ratio | 1:1 | All four came back 1024×1024 at 1K |
| Resolution | 1K to test the geography, 2K once the map is right | The map, not the model, is what gains from the extra pixels |
| Attempts | One per city, no retries, for everything in this post | Four of four were worth keeping as pictures; one, Lisbon, fails the shape check below and I would regenerate it |
| Landmark slots | Four, described by shape | Four was enough to make every city recognisable; I did not test how far you can push it |
| Negative clause | Always on | No run produced a visible logo with it in place |
The resolution choice is worth a note, because it is the opposite of what you might expect. The miniature city is small in the frame and looks fine at 1K. What should gain from 2K is the map: the place labels, the road shields, the contour texture. A label that is merely soft at 1K should sharpen; one that is already nonsense will just be nonsense in higher resolution. Everything in this post ran at 1K, so treat that as the reason to test at 2K rather than a measured result.
Both GPT Image 2.5 models run in the browser on the GPT Image 2.5 page without a ChatGPT account. An image costs 8 credits at the default medium quality and 1K; high quality and 2K or 4K cost more, and the generator shows the price before you run. New accounts start with free credits, which is enough to test a city or two before you commit to a series.
Five failures and their fixes
Five things can go wrong here. Two of them I hit; the rest are the ones worth having a fix ready for.
| What you see | Why | Fix |
|---|---|---|
| The model sits on the map like a game piece, with a visible seam | The "growing out of / emerging rather than sitting on top" pair got dropped | Restore both sentences; they are redundant on purpose |
| Whole frame is in focus and it looks like a video game | Depth-of-field line missing or weakened | Restore "shallow depth of field so the city is sharp and the surrounding map softly blurred" |
| The city is in the centre but the map beneath it is generic | "True geographic position" missing, or the region is too large to have a recognisable shape | Restore the phrase; name a region, not a continent |
| Coastline is wrong for the place | The model is drawing from memory of how that area looks | Regenerate rather than edit. If it fails twice, name the water body in the region slot, such as "the Tagus estuary" or "Faxaflói bay" |
| Labels are dense and half of them are nonsense | Too much map area relative to output size | Either raise resolution, or add fewer place labels, only major names |
The region slot is the one most people under-use. "Portugal" is a country; "the Tagus estuary and the coast around Lisbon" is a place with a shape. The narrower and more physically distinctive the region, the better the base map, and naming the actual water body helps more than naming the province.
Running a GPT Image 2.5 map series across cities
If you are making more than one, say a series for a travel account or a set of city cards, the variable to lock is not the city. It is everything else.
Keep one style block byte-identical across every prompt. The paragraph from "Realistic, refined, collector-grade model quality" to the end is the style block, and editing it between cities gives you a set that does not sit together. Change only the city, the region and the four landmark slots.
Keep the aspect ratio fixed too. 1:1 is the right default here because the diorama wants to sit in the middle with map bleeding off every edge, and square gives equal margin on all four sides. 3:2 should give you a wider band of map either side. I would avoid anything taller than square, since the extra height comes out of the map rather than the city and the map is what carries the effect.
You can also run the same city twice and pick. A second attempt is a cheap way to try for a better coastline without touching the prompt.
Which GPT Image 2.5 model to use for map dioramas
Everything here ran on Flare, the default, and for this genre I would leave it there.
The work in these images is texture, depth of field and a lot of small structure, which is what Flare already does well, and the failures that matter are geographic and textual rather than optical. Extra rendering time does not buy more geography. If you are producing a single final frame that someone will look at closely, or printing one large, our Flare and Sunburst comparison covers where the extra time earns its keep. For a run of ten cities, Flare is the sensible default.
Beyond cities: campuses, parks and nautical charts
The template is really "a real map with a three-dimensional thing growing out of the right spot", and a city is only one thing you can grow there.
Swap the city for a university campus, name the buildings by shape, and you have an orientation graphic. Swap it for a national park, with a ridge line, a lake and a visitor centre, and you have a trailhead illustration. Swap the topographic map for a nautical chart with depth soundings and put a harbour on it. Each keeps the style block and changes only the subject and the map type.
The one substitution worth avoiding is a real, current, named private property. The technique is good enough at architecture to produce something identifiable, and that carries a risk a pretty Kyoto does not.
Selling AI map diorama prints
People ask about selling these almost immediately.
Images you generate here can be used commercially, subject to the model's usage policies, and our commercial use guide covers the general position. Three things are specific to maps and worth thinking about before you list a print.
The place names are a factual claim on a product. A decorative poster with a misspelled city or a mountain at the wrong height is a different problem from a social post with the same error. A buyer can notice, and a refund request is harder to argue with than a comment. If you are selling, pass three is mandatory, and the shortcut of asking for an unlabelled map is worth considering: a diorama with beautiful nameless cartography has no wrong facts in it at all.
Print resolution is not the same decision as screen resolution. 1K is fine on a phone and thin on paper. For anything larger than a postcard, generate at 2K or 4K, and check the labels again at that size. A label that was passable when soft can become legibly wrong when sharp.
On a product, silhouettes matter more than they do on a social post. A recognisable building in a tiny diorama is normally fine. A building whose shape is itself a protected mark, rendered large on a print, is not.
A three-step GPT Image 2.5 map workflow
- Build one city and judge the map, not the model. Look at the coastline first. If the geography is wrong you have a prompt problem or a bad draw, and it is better to find that out before you fall in love with the miniature.
- Lock the style block before city two. The temptation is to keep improving the wording. That is exactly what breaks a set.
- Run the three checks before anything leaves your machine.
You can build all of it in the browser on the GPT Image 2.5 page with Flare already selected, or from the text-to-image generator if you want to change ratio and quality between attempts.
The Bottom Line
One prompt, six slots, and a style block you do not touch. The model will give you a convincing miniature almost every time and a trustworthy map only some of the time, and it will not tell you which one you got. Table Mountain's elevation was right to the metre in the same batch where Lisbon grew a peninsula that does not exist. Check the shape, read the labels, verify any number, and you can run the same prompt down a list of cities.
Related reading
- GPT Image 2.5 Flare vs Sunburst: 6 Same-Prompt Tests
- GPT Image 2 Style Library: 12 Copy-Paste Art Style Prompts
- GPT Image 2 Prompt Writing Guide: 7 Rules for 90% Hit Rate
- GPT Image 2.5 for Scientific Figures: What Journals Accept
gpt-image2.art is an independent site and is not affiliated with OpenAI. Place names and elevations in the generated maps were checked by hand where noted, and should be verified before any factual use.
Mais Publicações

10 Cinematic Camera Shot Prompts for GPT Image 2
10 copy-ready cinematic camera shot prompts for GPT Image 2, plus a practical shot formula, continuity workflow, and composition fixes.

Prompt Reverso do GPT Image 2: Reproduza Qualquer Imagem
Guia prático de prompt reverso para o GPT Image 2. Suba qualquer imagem de referência e obtenha um prompt reprodutível em segundos. 4 técnicas + templates prontos.

O Que é GPT Image 2? Uma Introdução Completa
GPT Image 2 é o modelo multimodal de imagem de próxima geração da OpenAI — o primeiro a lidar com texto não-latino e layouts complexos de forma confiável. Tudo o que você precisa saber.
Generate your first image with GPT Image 2 — right now
Reliable non-Latin text rendering, directed editing, and 50+ ready-to-use prompts. No downloads — just open in your browser.