Back to Blog

How to Write Prompts the Nano Banana Models Actually Follow

E2X Team··16 min read
nano-bananaprompt-engineeringimage-generationprompting-guidegoogle

Most prompting guides are lists of adjectives. Adjectives are the part of a prompt that matters least.

We run all four Nano Banana models behind one API, and the ways a prompt fails are consistent enough to write down. Forty words spent on quality adjectives and two on the light. A Pro-length compositional brief fired at the cheapest tier. An edit request that re-describes the entire scene and comes back as a different picture. None of these are model problems. They are writing problems, and writing problems have fixes.

A row of sharpened graphite pencils, a kneaded eraser and three overlapping sheets of blank cream tracing paper on a draughtsman's desk in soft window light

Medium, subject, light, frame — in that order

A prompt is read as a whole, but the early clauses do more work than the late ones. They set the space the rest of the sentence has to live in. Open with "a beautiful woman" and the model has to guess whether it is making a photograph, an oil painting, or a 3D render, and it will guess from whatever your other words statistically imply. Open with "35mm editorial photograph" and every following decision is already constrained to one medium.

So the order:

  1. Medium. The largest fork in the tree. Photograph, isometric render, flat vector illustration, gouache painting, macro product shot. Pick one and name it in the first four words.
  2. Subject. Not the category — the specific thing. "A woman at a café table" is a stock photo. "A woman in her fifties, both hands around a white cup" is a picture.
  3. Light. Direction, quality, colour. This one sentence moves the result more than any other, and almost nobody writes it.
  4. Frame. Camera height, distance, crop, depth of field, and where the empty space goes.

Here is the same request written badly and then written properly.

A beautiful woman in a coffee shop, highly detailed, 8k, masterpiece,
trending on artstation, ultra realistic, professional photography, award
winning, sharp focus, bokeh
35mm editorial photograph. A woman in her fifties sits alone at a marble
café table, both hands wrapped around a white cup, looking out of frame to
the left. Low winter sun through the window behind her, rimming her hair and
blowing the street outside to white. Shot from the next table at seated eye
level, waist up, shallow depth of field, background falling off to soft grey.

What changed: the medium moved to the front, the subject got one specific gesture, the light got a direction and a temperature, and the camera got a position. Nothing was added about quality. The second prompt is not more detailed in the vague sense — it is more decided.

And stop writing "highly detailed, 8k, masterpiece". It does nothing on these models. Those tokens are cargo from 2022 Stable Diffusion prompt culture, where they nudged a much smaller model toward a particular slice of its training data. Gemini-generation models do not respond to them, and every one of them is budget you could have spent describing where the light comes from.

Four sheets of blank translucent tracing paper stacked with slight offsets on pale linen, a single wooden pencil resting across the top sheet

The empty space gets filled unless you claim it

These models do not leave things out. If you describe a mug and say nothing about what surrounds it, the image will still contain a surround, and it will be whatever a mug photo usually contains: a wooden table, a linen napkin, a sprig of something, a soft-focus window, and roughly forty percent of the time, a word printed on the mug.

That is the single most common complaint we hear, and it is not a quality issue. Nobody said not to.

A ceramic mug on a table, product shot
Macro product photograph of a single unglazed stoneware mug, matte oatmeal
body, thumb-worn handle, standing dead centre on a seamless mid-grey sweep.
Large soft source from the upper left, one gentle shadow falling to the lower
right. No text on the mug, no logo, no other objects, no hands, no props, no
background detail.

The last sentence of that rewrite is doing more work than the first three combined. Write the exclusions as a plain list at the end of the prompt, in the same register as the rest of it. "No text, no logos, no people, no clutter" is often the highest-value line in the file — and on any image that will sit behind real typography later, it is not optional.

This matters more than usual for anything you plan to composite. A generated background with an invented sprig of rosemary in it is a background you have to mask.

A large expanse of empty warm-white paper with a single brass-ferruled drafting pencil at the lower right, soft diffuse light from above

Write for the tier you are actually calling

The four models take different amounts of prompt, and they want different shapes of prompt. The ceilings are not where people expect them:

ModelPrompt ceilingWhat it wants
Nano Banana 2 Lite32,000 charsShort, one subject, flat clause structure
Nano Banana 220,000 charsMedium length, holds two or three related elements
Nano Banana Pro50,000 charsLong compositional briefs with spatial relationships
Nano Banana (legacy)—Retires 2 October 2026; migrate rather than tune

Nano Banana 2 has the lowest ceiling of the three current tiers despite sitting in the middle on price. That surprises everyone the first time.

Ceilings are not the real constraint though. Lite is a lite model, and long prompts fail on it in a specific way: it keeps the first clause and the last clause and quietly drops the middle. Nested conditions ("a room with a table on which sits a bowl containing three pears, one of which is bruised") come back with the pears missing entirely. Write flat sentences for Lite. One subject, one setting, one light source, one exclusion list.

An isometric 3D render of a modern open-plan office at golden hour with
fourteen distinct employees at standing desks, a glass-walled meeting room on
the left containing a whiteboard, potted fiddle-leaf figs in the corners, a
coffee bar along the back wall with a barista, exposed ductwork above,
polished concrete floors reflecting the windows, and a cat asleep on the
third desk from the front
Isometric 3D render of an open-plan office at golden hour. Rows of standing
desks, a glass-walled meeting room on the left, tall windows throwing long
warm light across a polished concrete floor. Muted palette, clean geometry.
No text, no signage, no people.

The first prompt is not wrong. It is wrong for Lite. Send it to Pro and the fourteen employees, the ductwork and the cat all arrive. Send it to Lite and you get an office, a window, and a bill for $0.0238.

Which is the useful version of the advice going the other way: if your prompt is under sixty words with one subject in it, Pro buys you nothing except a thirty-second wait and roughly triple the cost. Most of what people generate is Lite work. We would rather tell you that than sell you the top tier for a placeholder background. For the full which-tier-for-which-job breakdown, we wrote that up separately in our four-way Nano Banana comparison.

Aspect ratio is a composition instruction, not a crop

Decide the shape before you write the sentence, because the shape changes what the sentence should say. A 21:9 banner wants a horizon and a subject pushed off-centre. A 9:16 story wants a vertical stack and headroom. If you write a prompt for a square and render it wide, you get a square composition with padding on either side, and it will look exactly like that.

Say the composition out loud in the prompt. "Horizon in the lower third, subject at the right third, empty sky across the upper left for copy" is a real instruction and these models follow it well.

Then there is the trap that costs people an entire asset batch. The default aspect ratio is not the same across tiers. On Nano Banana, Nano Banana 2 and Nano Banana Pro our default is 9:16. On Nano Banana 2 Lite it is 1:1. So a config that never pins aspect_ratio produces portrait images on three models and squares on the fourth, and if you switch tiers to save money the shape of your output silently changes underneath you.

Pin it. Always, on every call, even when the default happens to be what you want:

curl -X POST https://api.e2x.ai/v1/jobs/submit \
  -H "Authorization: Bearer $E2X_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "google/nano-banana-2-lite/text-to-image",
    "input": {
      "prompt": "Wide landscape photograph of a granite ridge under flat overcast light, horizon in the lower third, the ridge line entering from the right, empty pale sky across the upper left. No text, no people, no structures.",
      "aspect_ratio": "21:9"
    }
  }'

One more reason to know your tier: Lite carries four aspect ratios nothing else on our catalog has — 1:4, 4:1, 1:8 and 8:1. Those are genuinely useful and genuinely under-used. A site-wide header rail or a tall sidebar unit rendered at 8:1 beats generating a 16:9 and cropping 80% of it away, both in composition and in what you paid for the pixels you threw out.

Six empty wooden viewfinder mats and picture frames in different proportions arranged overlapping on a warm oatmeal paper backdrop

Editing is a different skill: write the delta

This is where the most experienced prompt writers regress hardest, because the instinct from text-to-image is to describe the picture you want. In an edit call you already have the picture. Describing it again competes with it.

An edit prompt has two jobs: name what changes, and name what must not. Nothing else belongs in there.

A woman in her fifties at a marble café table in low winter light, wearing a
red coat instead of a blue one, same face, same pose, editorial photograph,
35mm, shallow depth of field, warm tones
Change the coat from navy to deep brick red. Keep the fabric weight, collar
shape and drape identical. Everything else in the frame is unchanged: same
face, same expression, same hands, same cup, same window light, same crop.

The first version re-specifies the medium, the lens and the tone, and every one of those re-specifications is a fresh instruction the model may act on. That is how you get a subtly different face. The second version touches one thing and explicitly freezes the rest.

Two mechanical facts worth knowing before you build an editing pipeline. Reference counts differ by tier: the legacy model takes three images, Nano Banana 2's edit endpoint takes fourteen, and Nano Banana Pro's also takes fourteen but splits them by role — up to five character images for identity, six object images for product fidelity, three style references. That role split is the reason Pro can assemble a full ad composite in a single request, and it is a whole topic of its own; we cover it in keeping a character consistent across Nano Banana generations.

In-image text, and when not to bother

Legible words inside an image arrived with the Gemini 3 generation. The legacy gemini-2.5-flash-image cannot do it — a single word sometimes survives, a sentence never does. For anything with real typography in it you want Nano Banana 2, and for anything where the typography is the point you want Pro, which renders styled and multilingual text well enough that you can hand it an English infographic and get the Spanish version back with the artwork intact.

When you do ask for text, the rules tighten:

  • Put the words in quotation marks, exactly as they should appear.
  • Say how many words there are and where they sit in the frame.
  • Describe the letterforms as a style, not a font name. "Tall geometric Art Deco capitals, generous letter-spacing" works. Naming a specific typeface mostly does not.
  • Close with "no other text anywhere", or the model will add a plausible subtitle you did not ask for.
A vintage travel poster for Lisbon with the word LISBOA in art deco letters
Vintage 1930s travel poster illustration, flat screen-print style with
visible registration offsets. A cream tram climbing a steep pastel street,
the Tagus estuary behind it in three flat blues. One line of text only: the
word "LISBOA" in tall geometric Art Deco capitals across the lower quarter,
deep ochre on cream, generous letter-spacing. No other text anywhere in the
image.

Now the unpopular part. Most in-image text should not be in the image. If the words will ever need to change, be translated, be selected, be read by a screen reader, or be corrected after legal looks at them, generate the artwork clean and set the type in HTML, Figma or your template engine. Baking a headline into pixels because a model can do it is a decision you will pay for at the first copy revision. Ask for rendered text when the type is genuinely part of the artwork — a poster, a painted shopfront, a book spine in a still life — and skip it otherwise.

This has a sharper edge for anything data-shaped. A model that renders plausible text renders plausible numbers and plausible place names too, which is fine on a poster and catastrophic on a chart or a map. We went into exactly why in why AI image generators are so bad at maps.

Templates worth keeping

Four skeletons that cover most of what people actually generate. Fill the brackets, delete what does not apply, keep the exclusion line.

Clean subject on a plain ground, for compositing:

[Medium: macro product photograph / studio still life] of a single [subject
with one specific material detail], centred on a seamless [colour] sweep.
[Light source and size] from the [direction], one shadow falling [direction].
No text, no logo, no props, no hands, no background detail.

Editorial scene with copy space:

[Medium] of [subject doing one specific thing], [setting]. [Light: direction,
quality, colour temperature]. Shot from [camera height and distance],
[crop], [depth of field]. [Subject] at the [left/right] third, empty [surface
or sky] across the [region] for copy. No text, no logos.

An edit that changes one thing:

Change [element] from [current state] to [target state]. Keep [the two or
three properties of that element that must survive] identical. Everything
else in the frame is unchanged: same [list the parts most at risk].

Illustration with a single line of type, Pro tier:

[Illustration style and era], [technique detail]. [Scene described in two
sentences]. One line of text only: the word "[WORD]" in [letterform
description], [colour] on [colour], positioned [where]. No other text
anywhere in the image.

When prompting is not the problem

At some point the prompt is fine and the pipeline is the problem: retries, seeds, queueing, storing the output before the URL expires. That is a different discipline and we wrote it up in how to auto-generate images at volume. And if the gap you are trying to close is specifically photographic believability rather than instruction-following, the checklist for that lives in our guide to the most realistic AI image generator.

Every model publishes its exact parameter list — full aspect ratio set, defaults, what it refuses — as a machine-readable spec, like the Nano Banana 2 Lite spec file. Read that before you rewrite a prompt that started misbehaving after a model switch. Nine times out of ten the prompt was fine and a default moved underneath it. There is a spec file behind every model in the text-to-image category, and behind everything else we run.

Frequently asked questions

How should I structure a Nano Banana prompt?

Medium first, then subject, then light, then framing, then a list of exclusions. Naming the medium in the opening words constrains every decision that follows, and describing the light explicitly does more for the result than any amount of quality adjectives. Finish with what must not appear — "no text, no logos, no props" — because these models fill unspecified space rather than leaving it empty.

Do quality words like "8k" and "masterpiece" help on Nano Banana?

No. Those tokens came from earlier open-weight diffusion models where they steered output toward a particular subset of training data. Gemini-generation models do not respond to them, and they consume prompt budget that would be better spent on light direction, camera position or exclusions.

How long can a Nano Banana prompt be?

Nano Banana 2 Lite accepts up to 32,000 characters, Nano Banana 2 up to 20,000, and Nano Banana Pro up to 50,000. Those are hard ceilings, not targets. Lite in particular performs worse on long prompts well before it reaches its limit, because it tends to keep the opening and closing clauses and drop the middle.

Why do my images come out the wrong shape when I switch models?

Because the default aspect ratio differs by tier. Nano Banana, Nano Banana 2 and Nano Banana Pro default to 9:16 on our API, while Nano Banana 2 Lite defaults to 1:1. A request that never sets aspect_ratio will therefore change shape when you move between those models. Set it explicitly on every call.

How do I write an editing prompt for Nano Banana?

Write the delta, not the scene. Name the single thing that changes, name the properties of it that must survive, then state that everything else is unchanged and list the elements most at risk — face, pose, lighting, crop. Re-describing the medium, lens or mood in an edit prompt is what causes unintended changes elsewhere in the image.

Which Nano Banana model can render readable text in an image?

Nano Banana 2 and Nano Banana Pro can, because legible in-image text is a Gemini 3 generation capability. Nano Banana Pro is the reliable choice when typography is the point, since it also handles styled and multilingual text. The legacy gemini-2.5-flash-image cannot render legible text at all.

Do I need Nano Banana Pro for better prompt following?

Usually not. Pro's advantage is holding a long compositional brief with several spatial relationships, plus role-separated reference images and dependable in-image text. If your prompt is a single subject in under sixty words, Nano Banana 2 Lite at $0.0238 will follow it just as faithfully and return it faster.