Back to Blog

Consistent Image Generation With the Nano Banana API

E2X Team··13 min read
nano-bananaconsistencyreference-imagescharacter-consistencyimage-editingapi

Generating one good image is close to solved. Generating the twentieth that still belongs with the first nineteen is where projects die.

The gap is not prompt skill. It is that "consistency" is three different engineering problems wearing one word, and the technique that fixes one does nothing for the other two. We run the Nano Banana models on our API, and this is the request people get wrong most often — not because the models cannot do it, but because the references go in unlabelled and the model is left guessing what each one was for.

Twelve identical matte ceramic vases arranged in a four by three grid on a warm sand backdrop, each lit the same way and casting the same shadow

Three problems, not one

Separate these before you write a prompt. They fail differently and they cost differently.

Character consistency. The same person across many frames — same face, build and hair, recognisable as one individual whether seated in frame four or walking in frame nineteen. An identity problem: the model has to carry a face across compositions it has never seen.

Product fidelity. Not a similar bottle. That bottle. Exact cap proportion, label placement, shade of green. Fidelity is stricter than identity — a face has tolerance, a client's packaging does not.

Style consistency. The set reads as one set — same colour grade, same lens character, same light quality, same grain. No object needs to repeat; the treatment does.

Most briefs need two of these at once and ask for them as one instruction. "Make it consistent" is not a request an API can act on. "Image 1 is the person, image 2 is the bottle, image 3 is colour grading only" is.

The role split is the actual feature

Nano Banana Pro accepts up to 14 reference images. So does Nano Banana 2, at nearly half the price. The count is not the difference. Pro splits that budget by role:

RoleBudgetWhat it holds
Characterup to 5Identity references for a person or figure
Objectup to 6High-fidelity products, props, packaging
Styleup to 3Colour grade, lighting mood, treatment only

Fourteen total, but allocated. Allocation is what turns a reference image from a suggestion into an instruction.

Think about what happens without it. You hand a model fourteen pictures and a sentence, and it infers from the prose which picture is a face to preserve, which is a product to reproduce exactly, and which is only there for mood. It often infers wrong. The tell is a product that came out "in the style of" your packaging rather than as your packaging, or a face wearing the mood board's grade.

With the split you are not asking, you are declaring. Five slots for who, six for what, three for how it looks. One ad composite in a single request instead of three rounds of compositing.

Pro is $0.075 a request at our price on 26 August 2026, against a $0.15 list price. One more thing worth knowing: 1K and 2K cost the same on Pro. For hero assets, take the 2K. It is free.

An overhead view of a sorting tray with three compartments, one holding identical wooden blocks, one identical metal cylinders, one folded fabric colour swatches

Labelling references in the prompt

The role split does the structural work. The prompt does the semantic work, and it has to be explicit in a way that feels almost rude. References arrive in order, so name them by position and say what each is for, including what it is not for. That last part is where most prompts leak.

Image 1 and image 2 are the same woman — preserve her face, hair length and
build exactly. Image 3 is the product: reproduce the bottle shape, cap
proportion and label artwork precisely, do not restyle it. Image 4 is colour
grading and lighting mood only — do not copy any object, pose or composition
from it.

Place the woman from images 1-2 holding the bottle from image 3, seated at a
café table by a window, three-quarter view, warm afternoon light matching the
grade in image 4.

Four things make that prompt work, none of them a fifth adjective:

  • Every reference is accounted for. No image goes in without a stated job.
  • The style reference carries an explicit exclusion. Without it, a mood board leaks its furniture into your scene.
  • Identity words are concrete — face, hair length, build — not "looks like her".
  • The product instruction says "do not restyle", which is the phrase that separates fidelity from inspiration.

We are not going to cover general prompt construction here; that is the prompting guide. This is specifically about telling the model what your references are for.

The request

Consistency work is editing, not generation, so it goes to the edit-image slug with image_urls.

curl -X POST https://api.e2x.ai/v1/jobs/submit \
  -H "Authorization: Bearer $E2X_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "google/nano-banana-pro/edit-image",
    "input": {
      "prompt": "Image 1 and image 2 are the same woman - preserve her face, hair length and build exactly. Image 3 is the product: reproduce the bottle shape, cap proportion and label artwork precisely, do not restyle it. Image 4 is colour grading and lighting mood only, do not copy any object or composition from it. Place the woman holding the bottle, seated at a cafe table by a window, three-quarter view, warm afternoon light.",
      "image_urls": [
        "https://cdn.example.com/refs/model-face-front.jpg",
        "https://cdn.example.com/refs/model-face-three-quarter.jpg",
        "https://cdn.example.com/refs/bottle-front.jpg",
        "https://cdn.example.com/refs/grade-warm-afternoon.jpg"
      ],
      "aspect_ratio": "4:5",
      "resolution": "2K"
    }
  }'

Submit returns a job ID; poll it or pass webhookUrl. Pro takes around 30 seconds because it generates interim "thought images" while composing — Google does not return those and we do not bill them. Set aspect_ratio explicitly, since Pro defaults to 9:16. The full parameter walkthrough is in the Nano Banana Pro API tutorial.

Where the cheaper tiers land

We would rather talk you down a tier than sell you Pro for a job that does not need it.

Nano Banana 2 takes the same 14 references at $0.04, without the role split. That is genuinely fine for style consistency, where every reference is doing the same job and there is nothing to disambiguate. Six blog headers sharing a palette and a lighting mood do not need Pro. They need one good style reference and a frozen prompt template. Nano Banana 2 also reaches 4K, which Lite cannot touch at all.

Nano Banana 2 Lite is $0.0238 and built for volume, not fidelity. Reach for it when you need four hundred images that share a look and none of them carries a client's face or packaging. It is 1K only, with no resolution parameter — sending one is an error, not a no-op.

The legacy Nano Banana tops out at 3 reference images, which is not enough to do character and product in the same frame, and Google retires it on 2 October 2026 — the full story on that model is a separate post. Do not start a consistency project on it.

Rough rule: style-only work goes to Nano Banana 2 or Lite. Anything where a specific human face or a specific physical product has to survive intact goes to Pro. The price gap between $0.04 and $0.075 stops mattering the moment a client rejects a batch because the label was wrong.

Prices and product terms can change. Check each provider's current pricing before making a purchasing decision. Google Batch or Flex pricing is not directly equivalent to a standard on-demand API request because scheduling, availability, and processing conditions differ; it is therefore excluded from this comparison. This is a scoped comparison, not a claim that E2X is the world's cheapest option in every configuration.

Seed a canonical reference set once

The habit that pays for itself fastest here has nothing to do with the API.

Build the reference set once, deliberately, and reuse it forever. Not "whatever photos we had lying around" — a fixed, versioned, stored set every request draws from. For a character: several angles of the same face under neutral light, front, three-quarter, profile, plus one full-body for build. For a product: front, back, one detail crop of the label. For style: one or two images carrying the grade and nothing distracting.

Then freeze it. Store the URLs somewhere permanent — your own bucket, not a temporary link — and pass the same array in every request for the project's life.

This matters more than prompt tuning because drift is cumulative. Generate frame 5 from frame 4 and frame 6 from frame 5, and small errors compound until frame 20 is a different person. Anchoring every frame to the same canonical set means each generation drifts from the original by one step, never twenty. A fan, not a chain.

A related habit: if the set includes images you generated, keep the source files. Delivery URLs from any image API expire, so download and store your own assets — how to run generation automatically has the full loop.

Generate a hero, then derive the variants

The workflow that produces the most usable output per dollar is not "generate twenty images". It is "generate one image twenty times over".

Make one hero shot on Pro and get it right — composition, lighting, face, product. Burn a few attempts; at $0.075 the exploration is cheap.

Then stop generating and start editing. Feed the approved hero back in as a reference alongside your canonical set and ask for the change you want: different angle, different background, different crop, model looking away instead of at camera. Each variant is anchored to something that already passed review rather than to a fresh roll of the dice.

Editing beats regenerating for a reason that is easy to miss. A regeneration re-decides everything; every parameter you were happy with is back in play. An edit holds what you did not mention. Having spent an afternoon getting a face right, re-deciding it twenty times is the last thing you want.

Two practical notes. Watch the reference count as you add the hero back in — hero plus a five-image character set plus a product is already seven of fourteen. And save the approved hero's prompt next to the image; a variant is usually that prompt with two clauses swapped, and rewriting it from memory reintroduces the drift you just paid to remove.

One large pale ceramic vase in sharp focus with four smaller vases of the same form receding behind it, each turned to a slightly different angle

The part no reference set fixes

Google stamps SynthID into every image its models produce. It is invisible, it travels inside the file, and it is not a setting anyone exposes because it is not a setting at all. That matters here more than on most posts: consistency projects are client projects, and a client's contract may carry a clause about AI-generated assets. Settle that clause before you shoot the reference set, not after twenty approved frames are sitting in a deck.

Consistency work is nearly always editing, so the image-to-image models are the shelf to browse first, and the catalog has the video side if the brief moves.

Frequently asked questions

How do I keep the same character across multiple AI-generated images?

Build a canonical reference set — several angles of the same face under neutral light, plus one full-body shot for build — store it permanently, and pass the same images to every request. Use Nano Banana Pro's character slots, which hold up to 5 identity images, and state in the prompt which images are the person. Anchor every frame to the original set rather than to the previous frame, or errors compound into a different face by frame twenty.

How many reference images does the Nano Banana API accept?

Nano Banana Pro accepts 14, split by role: up to 5 character images, up to 6 object images, and up to 3 style references. Nano Banana 2 accepts 14 as well but without the role split. Nano Banana 2 Lite is built for volume rather than reference-heavy work, and the legacy Nano Banana caps at 3.

What is the difference between character consistency and product fidelity?

Character consistency means one person stays recognisable across many images — same face, hair and build, tolerant of small variation. Product fidelity means the object reproduces exactly: precise label artwork, cap proportion and colour, no restyling. Pro handles them in different reference slots because they need different strictness, and the prompt should say which images are which.

Should I use Nano Banana Pro or Nano Banana 2 for consistent images?

Use Nano Banana 2 at $0.04 when the need is stylistic — images sharing a palette, lighting and treatment, with no specific face or product to preserve. Use Nano Banana Pro at $0.075 when a particular face or physical product has to survive intact, because the role-split reference budget is what tells the model which is which. The price gap is small next to a rejected batch.

How do I label reference images in a Nano Banana prompt?

Refer to them by position in the order you pass them in image_urls, and state what each one is for. "Image 1 and 2 are the same woman, preserve her face and build; image 3 is the product, reproduce it exactly; image 4 is colour grading only, do not copy any object or composition from it." The exclusions matter as much as the instructions, since an unqualified style reference will leak its objects and composition into your scene.

Is it better to edit an existing image or generate a new one?

Edit, once you have an image you approve of. A regeneration re-decides every element, including the ones you were happy with; an edit holds everything you did not mention. The efficient workflow is one carefully made hero, then variants derived from it through the edit-image endpoint with your canonical references attached.

Why do my generated images drift away from the original character?

Almost always because each new image was generated from the previous one instead of from a fixed reference set. That makes errors cumulative — a slightly rounder jaw in frame 3 becomes the baseline for frame 4, and by frame 20 it is a different person. Anchor every request to the same stored canonical images so each output is one step from the original rather than twenty.