AI Images as Art References: Where They Help, Where They Lie
This question usually gets asked in bad faith by one side or the other, and we are not interested in either version. We run the image models and sell access to them, and we still think a working artist should refuse to use them for about half the jobs people currently use them for.
The split is not a matter of taste. It comes down to one distinction: a generated image is very good at being evocative and unreliable at being correct. Where your reference only has to suggest something — a mood, a colour relationship, the way light falls across a room at four in the afternoon — it works, and cheaply. Where it has to be right, where you are going to copy structure out of it, you get a confident, beautiful lie and you learn the lie.
Here is where we think the line sits, and how to actually get a useful reference out of the thing.

What generated references are genuinely good for
These all share a property: you take an impression from the image, not a measurement.
- Lighting studies. How does a warm key from a low window read against a cool bounce off a white wall. The model has absorbed a vast number of photographs of exactly this, and getting twelve variations of one lighting setup is trivial.
- Colour palettes. Pull the palette, throw the image away. Generated images are unusually good at coherent colour relationships because that is what the training data rewards.
- Composition thumbnails. Blocking, mass, where the eye enters and exits. A rough generated frame is faster to judge than a rough sketch.
- Mood boards. The most defensible use of all. Nobody is copying anything; you are pinning down a feeling before committing to a direction.
- Environment atmosphere. Fog density, dust in the air, wet asphalt at night, the green of a room lit through leaves. All impression, no structure.
- "What does this roughly look like from three-quarter view." Useful as a nudge when you are stuck, dangerous the moment you start tracing it. More on that below.

What they will quietly teach you wrong
The model does not know how a shoulder works. It knows what shoulders look like in photographs, which is a shallower kind of knowledge. Most of the time the difference does not show. When it does, it shows as a joint bending in a plausible direction it cannot actually bend in, and if you copy that into a drawing you have not made a mistake once — you have practised it.
That is the real cost, and it is why we are blunt about this list:
- Anatomy and joint mechanics. Especially anything foreshortened, anything under compression, anything where a muscle is doing work. This is exactly where a reference has to be correct.
- Hands and feet. Current models fail these far less than the 2024 generation did, but "far less" is not "reliably", and hands are the most-scrutinised thing in figure work.
- Perspective and architecture. Vanishing points drift, window mullions stop lining up across a facade, stair risers change height. It looks fine and it will not survive being built on.
- Machinery and how parts attach. A generated engine, a bicycle drivetrain, a suit of armour: the model renders the appearance of assembly without the assembly.
- Animal anatomy. Horses in particular: leg joints, gait phase, weight transfer. Photographers spent a century settling this; do not relearn it from a diffusion model.
- Period-specific costume, armour and tack. The model averages across centuries. A knowledgeable viewer spots it instantly, and then it is the only thing they remember.
The pattern: if you would have gone looking for a photograph of a real thing to check something, do that instead. Generated imagery is a substitute for a reference you were going to feel your way through, not for a reference you were going to measure.
The professional practice side
This matters more than people expect, and it is not mainly a legal question.
Provenance. Every image out of a Google model, all four Nano Banana tiers included, carries a SynthID watermark. It is invisible, embedded upstream by Google, and no provider can turn it off. Not us, not fal.ai, not Google's own API. Useful if you want provenance, an operational problem if a contract forbids watermarked assets. Know it is there before you build around it.
What you show a client. A generated reference belongs in your process, not your deliverable. Showing one at the concept stage sets an expectation for a picture you did not make and cannot reproduce on demand, and it is a quick way to be asked to "just use that one". If it is on a mood board, say what it is.
Why you do not trace it. Beyond the anatomy problem, tracing is where provenance gets sharp. A composition you arrived at by looking at a generated frame is yours. A shape lifted line-for-line off one has an ancestry you cannot document. Use it the way you would use a photograph you found and cannot license: look, understand, close the tab, draw.
None of which makes a stock photo purer. A stock sunset is also somebody else's decisions about light. The distinction here is correctness, not virtue.

Actually getting a useful reference
A reference does not have to be beautiful. It has to be legible. That changes which model you should use, and the answer is the cheap one.
Nano Banana 2 Lite costs $0.0238 an image on our API as of our 26 August 2026 price check. Twenty variations is $0.476, and nothing else in our catalog price breakdown undercuts it. That is the right economics for reference work: quantity and spread, not one perfect frame. Paying $0.075 for a Nano Banana Pro render you are going to squint at for ten seconds and discard is money set on fire. Save Pro for output, not for input.
Prices and product terms can change. Check each provider's current pricing before making a purchasing decision. Google Batch or Flex pricing is not directly equivalent to a standard on-demand API request because scheduling, availability, and processing conditions differ; it is therefore excluded from this comparison. This is a scoped comparison, not a claim that E2X is the world's cheapest option in every configuration.
Four habits get more out of it than better prompting does.
Generate the same scene at several aspect ratios. Framing changes what the model puts in the picture, not just how it crops. Lite carries fourteen fixed ratios plus auto, including 1:4, 4:1, 1:8 and 8:1 — those extreme ones do not exist on the other tiers and they are unexpectedly good for figuring out a tall composition or a long panel. One prompt, six ratios, six genuinely different thoughts about the same scene.
Iterate lighting on a fixed subject with the edit endpoint. This is the big one. Generate a subject you like, then push it through the Lite edit endpoint repeatedly, changing only the lighting: hard sun from camera left, overcast, one warm practical lamp set low, backlit at dusk. That is one subject under six conditions, which is what a lighting study actually is. For identity that has to hold across a longer series, Nano Banana 2 takes up to 14 reference images — covered in keeping characters consistent across generations.
Prompt for the light, not the subject. "Single hard source from high camera left, deep unfilled shadows, warm key against cool ambient" is the useful half of the prompt. What is standing in the light matters less than you think.
Keep the failures. A frame with a wrong hand in it is still a usable lighting reference. Binning everything with an error in it throws away most of what you paid for.
One warning that saves an afternoon: the output URL our API returns is temporary. Build a reference library by downloading the bytes, not by saving links. A folder of expired links is not a library.
Publishing a generated image instead of drawing from one is a stricter job, and we wrote the inspection habits for it separately: how to tell a generated photograph from a real one. Reference work really only needs two endpoints, text-to-image to make the subject and image-to-image to keep pushing it around, both listed with their parameters in the catalog.
Frequently asked questions
Is it effective to use AI images as art references?
For lighting, colour, composition, mood and environmental atmosphere, yes — those are impressions, and generated images carry impressions well. For anatomy, perspective, architecture, machinery and animal structure, no: the model reproduces what those things look like in photographs without knowing how they work, so copying its output means practising its errors. The test is whether your reference needs to be evocative or needs to be correct.
Can I trace over an AI-generated image?
We would advise against it. Two reasons: anatomical and structural errors get copied straight into your work when you trace rather than interpret, and a traced shape has an ancestry you cannot document if a client or a jury asks later. Use a generated frame the way you would use an unlicensed photograph you found online — look at it, understand it, then draw from understanding.
Do AI-generated reference images have a watermark?
Every output from a Google image model, including all four Nano Banana tiers, carries SynthID, an invisible provenance marker applied by Google. No provider can disable it, since it is embedded before resale. It does not affect the image visually and will not interfere with using it as a reference, but it does mean the file is identifiable as AI-generated.
What is the cheapest way to generate reference images?
Nano Banana 2 Lite at $0.0238 per image on E2X, checked 26 August 2026. Reference work rewards volume over polish, so twenty variations at roughly $0.48 beats one premium render. There is no reason to pay top-tier rates for something you will glance at and throw away.
How do I generate the same subject under different lighting?
Generate the subject once, then send that image to an edit endpoint repeatedly, changing only the lighting instruction each time. That gives you one subject under many conditions rather than several unrelated images, which is what makes it a study. Nano Banana 2 Lite handles this at $0.0238 per pass; Nano Banana 2 takes up to 14 reference images if you need tighter identity consistency.
Should I show generated references to a client?
Keep them in your process rather than your deliverable. A generated image on a client-facing board sets an expectation for a picture you did not make and cannot reproduce on demand, and clients routinely ask for "that one" once they have seen it. If a generated frame does appear on a mood board, label it.
Why do AI models get hands and anatomy wrong?
Because an image model learns the appearance of a hand across millions of photographs, not the structure underneath it. There is no skeleton in the representation, only statistics about how fingers tend to sit next to each other, so a plausible arrangement can be mechanically impossible. That shows most in hands and foreshortened limbs, which is exactly where artists need a reference to be accurate.