GROX
Technique

How do you describe an image to an AI so you get what you pictured?

Published 20 September 2026

Write the subject first, then the setting, then how it is framed, then the light, then the style. That order mirrors how a photographer briefs a shoot, and it gives the model the most important information before it has to guess. Everything else — aspect ratio, what must stay consistent, tone — is a layer on top of that foundation.

What does a well-structured image prompt actually contain?

Think of a prompt as five concentric rings. The innermost is the subject: who or what the image is about, described with enough specificity that a stranger could pick it out of a crowd. 'A woman' is a subject. 'A woman in her sixties, short grey hair, reading glasses pushed up on her forehead' is a subject the model can hold steady.

The next ring is the setting — where the subject exists. Indoors or outdoors, the time of day, the season, the mood of the place. Then framing: close-up, wide shot, bird's-eye, over-the-shoulder. Then light: the direction it comes from, whether it is hard or soft, warm or cool. Finally, style: photorealistic, flat illustration, oil painting, line art. Skipping any ring forces the model to guess, and its default guess is rarely yours.

Subject
The person, object or creature the image is about. Be specific enough that the model cannot substitute something plausible but wrong.
Setting
Where the subject exists — location, time of day, season, atmosphere. This shapes colour, shadow and mood before any style instruction does.
Framing
The implied camera position: close-up, wide shot, aerial, eye-level. It tells the model how much of the world around the subject to show.
Style
The visual language: photorealistic, watercolour, flat vector, pencil sketch. Without it, the model picks its statistical average, which is usually a smooth, slightly over-lit render.

Why does aspect ratio matter more than most people expect?

The default output shape for most image tools is square. A square is fine for a profile picture and wrong for almost everything else. A hero banner needs a wide rectangle. A phone wallpaper needs a tall one. A product shot for a marketplace thumbnail needs a square. If you do not say, you get a square, and cropping afterwards rarely recovers the composition the model built for the centre of the frame.

State the shape early — before style, ideally after framing — because the model uses it to decide where to place the subject. A portrait-format prompt with a centred subject will leave empty sky above and cut feet below if the output is forced square later. The shape is not a technical afterthought; it is part of the composition.

Common use cases and the aspect ratio that fits each
Use caseRatioNotes
Social square post1:1Safe default for Instagram grid and marketplace thumbnails
Landscape hero banner16:9Website headers, YouTube thumbnails, presentation slides
Portrait phone wallpaper9:16Stories, Reels, TikTok covers
Print A4 / letter3:4 approx.Posters, documents, editorial illustration
Cinematic wide21:9Film stills, panoramic backgrounds, desktop wallpapers

How do you keep elements consistent across a set of images?

A single image is forgiving. A set of images — a character appearing across six scenes, a product shown in three settings, a brand illustration series — demands consistency that a prompt alone cannot guarantee unless you are deliberate about it.

Name what must stay fixed and put it in every prompt in the set, word for word. If the character has a red scarf, write 'red wool scarf' in every prompt, not 'scarf' in one and nothing in another. If the lighting is overcast daylight in the first image, write 'overcast daylight, soft shadows, no direct sun' in every subsequent one. The model has no memory of what it generated a moment ago; each prompt is a fresh start. The consistency has to live in the text, not in your expectation.

  • Copy the fixed-element phrase exactly between prompts — paraphrasing introduces drift.
  • Describe what is absent as well as what is present: 'no hard shadows' is more reliable than hoping 'soft light' implies it.
  • If a character has a distinguishing feature, name it in physical terms rather than a proper name the model has no reference for.
  • Keep a short reference block — three or four phrases — that you paste into every prompt in the set before adding the scene-specific detail.

Why should you read the prompt the model actually used?

Many image tools show you the prompt that was submitted after any automatic expansion or rewriting. That text is more useful than the image itself when the result is wrong. If you asked for 'a quiet harbour at dusk' and the model expanded that to include 'fishing boats, orange sky, dramatic clouds, long exposure effect', you now know exactly which words to remove or replace — rather than guessing why the image feels busier than you wanted.

Reading the actual prompt also reveals the model's defaults: the adjectives it adds when you leave a gap. Over a few attempts you build a picture of what it assumes when you say nothing about light, or nothing about people. That knowledge makes your next prompt shorter and more accurate, because you only need to override the defaults you disagree with.

Common questions

How long should an image prompt be?

Long enough to cover the five rings — subject, setting, framing, light, style — and short enough that no ring is buried. Somewhere between thirty and eighty words covers most cases. Longer prompts are not more powerful; they are more likely to have one phrase cancel another. If you find yourself writing more than a hundred words, split the prompt into a core description and a separate list of exclusions.

What is the most common reason an AI image does not match what I pictured?

A missing or vague subject. 'A cosy room' gives the model almost nothing to anchor; 'a small living room with a worn leather armchair, a single floor lamp, and a stack of paperback books on the floor' gives it a great deal. The model fills every gap with its statistical average, which tends toward the generic. Specificity in the subject reduces the surface area for unwanted defaults.

Should I describe what I do not want in the prompt?

Yes, but sparingly. Negative instructions — 'no text', 'no people', 'no lens flare' — are useful when the model keeps adding something you have not asked for. They are less useful as a substitute for a clear positive description. Start with what you do want; add negatives only for elements the model inserts repeatedly despite their absence from your prompt.

Does the order of words in a prompt change the result?

Often, yes. Most image models weight earlier tokens more heavily, so what you write first tends to dominate the output. That is why subject comes first: it is the thing you most need the model to get right. Style instructions placed at the very start can overwhelm the subject; placed at the end, they modify it. If a style is taking over the image, move it later in the prompt.

If you want to go from a prompt to a finished asset — image, document, or deployed page — without switching tools, GROX handles creation, iteration and publishing from one conversational surface.