Generating images with AI
AI image generation means a model builds an image from your text description, one that did not exist before. The hard part is not the generating; it is getting that image to something you can hand a client.
- Lesson 11 of 12
- Beginner
- Free, no signup
From one sentence to a file you can hand over
-
1
The idea
What this image has to do and for whom. One sentence. Until you have it, no output can be judged.
-
2
Brief and prompt
Six explicit constraints: subject, action, frame, light, medium, palette. Leave one out and the model picks it.
-
3
Generate
The only step the model performs. Running the same prompt again gives a different image, not a better one.
-
4
Select and lock the look
Keep the right output as a reference input and build the rest of the set from it. This is where a set starts to match.
-
5
Edit
Crop, final colour, and any text meant to sit on the image. Persian text is added here, not inside the model.
-
6
Deliver
The right format, dimensions and weight, plus a decision about whether the client is told AI was used.
Step three is the only step the model performs. The other five are your work, and that is where the difference between a fun picture and a deliverable file is made.
Last checked: Facts and tool names in this lesson are re-checked against their sources on this date.
What an image model actually does
A model trained on millions of images and captions builds a new image from your sentence. It does not search, it does not pull a file from anywhere, and it does not find an image; every output pixel is made at that moment.
That distinction has one practical consequence most tutorials skip. Run the same prompt twice and you get two different images, not a better one. So the common habit, hitting generate until one of the outputs comes out well, is buying lottery tickets. The real lever sits somewhere else.
That lever is the reference input: one or more images you send alongside the prompt so the model takes shape, style and character from them. It is the only way a product or a person stays the same across ten images in a row, and it is what fills the gap between a fun picture and a deliverable set. The ceiling on reference images differs by model, and the numbers are in the next section.
One more thing that is invisible on day one: these models do two separate jobs and they are not equally good at both. Generating an image from text is one job; editing an existing image is another, and the model order differs between them. If your work is mostly editing, the text-to-image ranking is not the answer to your question.
And the last thing worth saying up front: the output is a frame, not a delivery file. Cropping, format, weight, final colour and any text that has to sit on the image are still your job. An image that goes straight from the model to a client site is usually both heavier than it should be and not quite what was asked for.
Which models actually have numbers today
The AI section of this site already ranks six image models with published weights and sources, and we are not rebuilding that table here; the image model ranking lives where it lives. What a designer should take from it is three lines.
Line one: GPT Image 2 is today the only model that wins both measured boards, text-to-image and editing. Line two: Nano Banana Pro, the expensive Google model, sits below Google own cheap model Nano Banana 2 and costs twice as much. Line three: of those six, exactly one publishes open weights you can run on your own hardware.
We re-read the prices from each vendor own page on the day this lesson was written, and all five matched what the AI section already records. The cheapest is FLUX.2 klein 4B at $0.014 per image. GPT Image 2 is $0.053 at medium quality and $0.211 at high, four times as much. Nano Banana 2 is $0.067 for a 1K image, FLUX.2 max starts at $0.07, and Nano Banana Pro is $0.134.
The reference-input ceiling has its own published numbers, and for repeat work it matters more than the rank: Nano Banana 2 accepts up to ten object images, Nano Banana Pro six, FLUX.2 up to eight on the API and klein 4B up to four. If one fixed product has to recur across a set, that column decides for you, not the quality column.
And one thing you only learn by running it: the row that came first on a leaderboard names a configuration, not a model. The first GPT Image 2 row is medium quality, the $0.053 one. Pick high quality and your cost is four times larger with no measured score behind that choice at all.
Why Midjourney appears in none of our tables
Midjourney is everywhere in Persian recommendation lists and nowhere in our table. Not because it is bad. Because we have read nothing of it.
We tried again on the day this lesson was written: the home page, the documentation and even the robots.txt file of the Midjourney site all return 403 to our server. So we cannot read its pricing page, we cannot read its terms of use, and we cannot even read the rule we are meant to obey. For comparison, the documentation sites of OpenAI, Google and Black Forest Labs all open from this same server.
And the part that matters more: Midjourney has no row on either measured image board. The AI section of this site last read those two boards on 12 August 2026: seventy-seven models scored on text-to-image and fifty-three on editing, and its name was on neither. So we have no facts and no evidence, and we do not build a row without numbers.
Our position is that plain: if you read somewhere that a tool is the best, ask where the number is. A list of ten tools with no measured column is SEO content, not a buying guide. This is not an argument against Midjourney; it is an argument against any recommendation without a source, ours included on the day we stop citing one.
How to write an image prompt
A good prompt is not a description, it is a brief. Instead of writing one poetic sentence, you fix six things explicitly and leave the rest to the model: subject, what it is doing, frame and camera distance, light, medium or style, and colour palette. Leave any of those six out and the model picks one for you, and you only find out which when you look at the output.
The part rarely said out loud: piling on adjectives usually does nothing. Ultra high quality, 8K, cinematic, award winning is a row of words that adds no real constraint. A real number or ratio does far more work: landscape frame, subject in the left third, empty space at the top of the frame for text. Those are the things you discover were necessary the moment you drop the image into a layout.
Write negative constraints too, but keep them short. A long list of what you do not want takes as long as the list of what you do want, and its effect is not predictable in advance. Two or three things that genuinely go wrong every time is enough.
And the last step, the one that makes the actual difference: when an output finally has the right look, keep it as a reference input and build the rest of the set from it. That is what the reference column in the previous section exists for, and it is the only way ten images in a set come out looking related. Anyone who does not know this is rolling the dice from zero every time.
The image prompt sheet
What the prompt must carry
- The subject and what it is doing
- Frame and camera distance, landscape or portrait
- Light: where it comes from and how hard it is
- Medium or style: photo, illustration, 3D render
- Colour palette, ideally with the brand codes
- The empty space you need for text
What wastes a generation
- A row of adjectives: high quality, cinematic, award winning
- Persian text you expect the model to write
- Re-running the same prompt hoping for a better result
- A long list of negative constraints
This sheet does not guarantee a good output; it only stops the model from deciding on your behalf and you finding out afterwards.
Persian text inside an image, where this breaks
For an Iranian designer this is the single most important question, and the honest answer is that no vendor has published a number about it.
Google describes text in its models as advanced. OpenAI writes that output may still have trouble with text. Black Forest Labs calls one of its variants specialised for typography. Three adjectives, zero measurements, and none of the three says anything about Persian at all. We have not tested these six models on Persian either, and until we do we will not print a number.
There is one practical recommendation that needs no number, and it comes from years of putting images on Persian sites: do not generate the text inside the image. Build the image with no lettering at all and put the text on top in your design tool or directly in CSS. You win three things at once: correct letterforms, the ability to translate that same image into another language, and text a search engine and a screen reader can actually read.
There are cases where this does not work, such as text that has to sit on a curved surface or inside the scene. Even there the answer is not to hope the model gets it right; the answer is an editing pass in a design tool.
Who owns the image you just made
For personal work this question is not serious. For work that has an invoice attached, the answer has to be known before the image is made, not after it is delivered.
In the additional terms for the Gemini API, Google states plainly that it will not claim ownership over content you generate, and then immediately adds the second sentence: you acknowledge that Google may generate the same or similar content for others and reserves all rights to do so. What you get is ownership, not exclusivity. For a logo that distinction is the whole story.
Black Forest Labs states its licences more clearly, because it publishes model weights: FLUX.2 klein 4B is Apache 2.0, so commercial work is free and it runs on a GPU with roughly thirteen gigabytes of memory. The 9B variant of the same family is under their own non-commercial licence, and the dev variant is non-commercial too. One family name, three different terms; if you read somewhere that FLUX is free, ask which FLUX.
About OpenAI and Midjourney we owe you a limit of our own: the terms-of-use pages of both return 403 to our server. OpenAI technical documentation does open and we quote only that; we write nothing about what we have not read. If you made an image with either and it is going into client work, read the terms from your own account.
And one point no licence solves: owning something is not the same as it being original. An image a model made can legally be yours and still resemble two hundred others produced with the same model that week. For a brand identity, that is a risk in itself.
Where to use these images and where not to
Our position is simple: an AI image does not replace a photograph, it replaces a stock photograph.
Anywhere the image is decoration and nobody makes a decision from it, these models beat any stock library on both speed and cost: a first draft of an idea for a client, a slide background, an explanatory shape inside an article, a mood board. Speed is worth something there and originality is worth nothing.
Anywhere the image makes a claim of its own, no. The product shot a customer buys from, the team photo, the portfolio piece, any image that says this is us or this is your product. A fabricated image in those places is not only a legal problem; it is a trust gap that does not close once somebody notices.
There is a third layer most tutorials never mention: what these models produce is usually not unmarked. Google image models put an invisible watermark on every output, and some services attach a signed record to the file as well. If you want to know what is inside the file you are uploading, the lesson on AI image watermarks and SynthID opens exactly that.
And if the output has to be a publish-ready file rather than a raw frame, that is the work done in our graphic design service.
The fast path, with AI
The fast path here is not a longer prompt. It is two moves that Persian tutorials almost never show: write the brief once and let a text model turn it into three genuinely different prompts, then instead of re-rolling, lock the right output and pull the rest of the set from it. The second move is what turns a single picture into a set. A fast, cheap model is enough to expand the brief; our current pick lives in <a class="text-link" href="/en/ai/">the AI section</a>, and take the image model from <a class="text-link" href="/en/ai/image/">that ranking</a>.
- In one sentence, write what this image has to do, for whom, and where it will sit. State the dimensions and the space for text here, not later.
- Give the recipe below to a text model with your own variables and take three prompts. Do not drop parts three and four; they are what pulls this out of guesswork.
- Run all three prompts once and pick one. If none is right, change the brief rather than the prompt: almost always the problem is in that first sentence.
- Feed the chosen image back as a reference input and build the rest of the set from it. Every new image is now a version of that look, not a fresh lottery.
- At the end, check two things yourself: the text on the image was added outside the model, and the licence of the model you used matches what the image is for.
Copy-ready recipe
Role: art director writing an image brief for the web.
What this image has to do and for whom:
{write it in one sentence}
Where it sits: {e.g. service page header, 1600 by 900, text over the left half}
Brand palette: {colour codes}
Never in this image: {two or three things}
Return exactly these four parts:
1. Three English prompts, one paragraph each, each stating all six constraints
explicitly: subject, action, frame and camera distance, light, medium or
style, colour palette. The three must differ in medium or style, not in
adjectives.
2. For each prompt, say which of those six constraints I did not give and you
filled in yourself.
3. At most three negative constraints, only things that genuinely go wrong on
this particular subject.
4. One sentence: after I pick one of the three and feed it back as a reference
image, which property of it should I lock so the rest of the set matches.
Put no text inside the image.
Before you trust the output: This path does not solve three things, and all three are a person job. First, text on the image: the prompt said no text inside the image, so writing it is still yours. Second, licensing: locking a look does not make the image original, and if the output is for a brand identity, Google sentence about generating similar content for others still stands. Third, what is inside the file itself: output from these models usually carries a watermark or a signed record, and you should know which before you publish.
AI in this kind of work
This whole lesson is about working with AI, so this section takes the question usually asked last: which of these is reachable from Iran at all, and what does the output carry with it. Our position in one sentence: <strong>a model whose weights you can download is more defensible for client work than a better model that only exists behind a foreign account, even at lower quality.</strong>
Tools that actually help
- GPT Image 2 Today the only model that wins both measured boards. At medium quality it is $0.053 per image, and that is the configuration that earned the rank. OpenAI does not publish a reference-image ceiling in its documentation, and Iran is not on its supported-countries list, so we know of no official payment route.
- Nano Banana 2 The cheapest route to a result close to first place, at $0.067 for a 1K image with a ceiling of ten reference images, the highest in this table. Google published list of available regions for the Gemini API does not include Iran. And one thing that matters in the next section: Google own documentation says every generated image includes a SynthID watermark.
- FLUX.2 The only family here with open weights: klein 4B is published under Apache 2.0, runs on a GPU with roughly thirteen gigabytes of memory, and commercial use is free. If your condition is that the image never leaves your machine and no account or card is needed, this is the only option. Its quality is behind the top of the table, and Black Forest Labs publishes no supported-countries list either.
- Nano Banana Pro It is here because it is the common mis-purchase: Google flagship at $0.134, twice its own cheap model, and ranked below it. If you are choosing between the two Google models, the cheaper one is almost always the answer. Like every Google image output, this one carries a SynthID watermark.
Where it backfires
Two specific risks in this work, and neither is about quality.
First, what leaves with the file. The Gemini API documentation states plainly that all generated images include a SynthID watermark, an invisible mark Google itself can detect. This is neither illegal nor secret; but it does mean a site publishing dozens of such images is identifiable from the files themselves. If you want to know exactly what is inside the file and who can read it, the next lesson in this track opens that.
Second, what a licence does not solve. Google states in its own terms that it does not claim ownership of the output, then immediately says it may generate the same or similar content for others. For a decorative image that sentence is unimportant; for a logo or a brand identity it is the whole story.
And on access from Iran we suggest no way around anything. To see where each of these tools stands, read AI in Iran and the buying guide.
Sources: Google: image generation in the Gemini API Google: Gemini API pricing, per-image tiers Google: Gemini API Additional Terms, Use of Generated Content OpenAI: image generation guide, cost per image Black Forest Labs: FLUX.2 overview and licences Black Forest Labs: pricing per image
Where this advice stops
This lesson did not run the six models side by side on one prompt; every number in it is published, so what you are reading is a comparison of documentation, not our test. Four things are deliberately absent. Any number about the quality of Persian text inside an image, because no vendor publishes one and we have not measured it. Anything from the terms of use of OpenAI or Midjourney, because both pages return 403 to our server and we do not write about what we have not read. Any number about how much time this workflow saves, because we have no source we would defend. And any route to access from Iran, because we do not teach what we do not recommend. One more honest boundary: the prices here are the cheapest published tier of each model, and a tier is not a model; the same model at higher quality can cost four times as much.
From our own work
That we do not build a row without numbers is a decision, and it is checkable on this very server. On the day this lesson was written we tried again: the Midjourney home page, its documentation and even its robots.txt file all returned 403 to our server. In the same minute, the documentation of OpenAI, Google and Black Forest Labs all opened with a 200, so the problem was not our network. You can see the consequence in the data of the AI section of this site: the Midjourney file exists, its status is live, and it deliberately contains no facts and gets no page. You can run the same test on any tool somebody recommended to you, and it takes a minute: open its pricing page and its terms page and see whether they open at all.
Real follow-up questions
Which AI is best for generating images?
If budget is not the constraint, GPT Image 2, because it is the only model that wins both measured boards. If cost matters, Nano Banana 2 gets close for less, and if your condition is that the model runs on your own hardware, FLUX.2 klein 4B is the only option. Three different answers for three different conditions, and none of them is best without a condition.
Can an AI image be used on a client site?
It depends what the image claims. For decoration, backgrounds and explanatory shapes, yes, and usually better than a stock library. For product shots, team photos and portfolio pieces, no, because there the image is itself a claim. And before any decision, read the licence of the model you used; across these five models there are three different sets of terms.
Can these models be used from Iran?
For GPT Image 2 and both Google models, Iran is not on the supported-countries list and there is no official payment route. For FLUX no country list has been published at all. The only path that needs no foreign account or card is running klein 4B locally. We do not sell accounts and we do not suggest ways around anything.