SEO

SEO for AI engines (GEO)

GEO is the work of getting AI assistants to point at your site in their answers, and today only one part of it is measurable: whether those services' crawlers reach your pages, which you can see in your own server log. The rest is either the vendors' own documentation or inference, and anybody handing you a list of GEO ranking factors is selling a guess.

  • Lesson 15 of 15
  • Advanced
  • Free, no signup

Where a citation in an assistant's answer comes from

Three inputs that reach the core, and one that deliberately does not connect, because Google itself writes that Search does not use it.

  • 1Crawler access

    The only thing you measure yourself: in the server log, with the robot name and the response code.

  • 2A quotable answer in the first two sentences

    Inference, not law. It is independently right for a human reader too.

  • 3Sourced claims and stable terminology

    Inference. If the model wants to check, the path should be right there.

  • A special file for AI

    Google writes that Search does not use these files and that they neither harm nor help. This wire does not reach the core.

  • The answer the assistant builds

    None of these companies has published its criteria for choosing a source. The burden of proof is on whoever claims otherwise.

  • A citation of your page

    And a citation is not necessarily a click. No tool hands you a count of them either.

Only the first input is measurable. The two after it are inference, and this figure does not turn them into law either.

Last checked: Facts and tool names in this lesson are re-checked against their sources on this date.

What is GEO, and which part of it is actually proven?

GEO is short for generative engine optimization, the name for work aimed at being visible inside the generated answers of search engines and assistants. AEO is the same thing under another name.

Google has an official position on both words, and reading it first will save you time: in its guide to optimizing for generative AI features it writes that from Google Search's perspective, optimizing for generative AI search is optimizing for the search experience and thus still SEO, and it adds that if you are considering third-party AEO or GEO advice or services, you should review its guidance on evaluating third-party SEO advice. We unpacked that position in detail in Google said not to build a special file for AI, and we will not repeat it here.

But Google is not the only engine in the world, and this lesson is about exactly the part that post does not cover. ChatGPT, Claude and Perplexity each run their own crawlers and their own policies, and each publishes its own documentation. There are things in there that are documented and certain, and almost nobody in the Iranian market talks about them.

So before any advice, let us drop the claims into four buckets and keep those four buckets to the end of the lesson: what you measure yourself, what the vendor documented, what is a reasonable inference, and what is pure guesswork sold as a package. Every sentence in this lesson sits in one of the four, and we say which one each time.

The four buckets every GEO claim falls into

Before accepting any recommendation, ask which bucket it is in. The first three are usable; the fourth was built to be sold.

  • Measured

    Visible in your own server log: which robot came, what it fetched, what code it got.

  • Vendor documented

    Written on the company's own official page and can be linked.

  • Reasonable inference

    It has an argument but no document. Do it, but do not sell it as law.

  • Guesswork being sold

    Lists of GEO ranking factors and guaranteed placement in answers. It has no reference.

This claim

The boundary between the second and third bucket is where most of the error happens: a correct recommendation gets sold as a Google rule.

Which crawlers come to your site, and what does each one do?

This section is the documented part of the lesson and the most valuable one, because it is the only lever actually in your hands and its outcome is binary: either they are allowed in or they are not.

The common mistake is to assume each company runs one crawler. They do not. OpenAI writes in its own documentation that OAI-SearchBot and GPTBot are two independent settings: one for being visible in search, the other for model training. A site that blocks GPTBot and leaves OAI-SearchBot open is out of training and present in search. The same page writes that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links, and that it can take around twenty-four hours from a robots.txt change for their systems to adjust.

Anthropic runs three separate robots and explains all three on its own support page: ClaudeBot, which collects web content that could contribute to model training; Claude-User, for when an individual asks a question and the system goes and fetches your site; and Claude-SearchBot, which works on the quality of search results. That page writes that disabling Claude-User prevents their system from retrieving your content in response to a user query, which may reduce your site's visibility.

And then the Google case, which is misunderstood more than any other. Google-Extended has no separate user agent string at all; Google's own documentation says crawling is done with existing Google user agent strings and that this token acts purely as a control in robots.txt. However long you search your log, you will never find a request calling itself Google-Extended. What the token controls is whether your content may be used for training future generations of the Gemini models and for grounding in those same products, and the same page states plainly that Google-Extended does not impact a site's inclusion in Google Search and is not used as a ranking signal.

robots.txt tokenwhat it controlswhat blocking it costs
GPTBotuse of your content for training OpenAI's foundation modelsyour content stays out of training
OAI-SearchBotappearing in ChatGPT's search answersyou are not shown in search answers, though you may remain a navigational link
ClaudeBotcollecting content for Anthropic's model trainingyour future materials are excluded from the training datasets
Claude-SearchBot and Claude-Usersearch indexing, and fetching a page in answer to a user questionyour visibility and accuracy in results and in user answers drop
Google-Extendedtraining and grounding of the Gemini modelsno effect at all on presence or ranking in Google Search

Read the table once from left to right and notice how simple the decision becomes. If you want to be visible in ChatGPT answers but keep your content out of training, you know exactly which two lines to write. The documentation of all three companies is in this lesson's source list.

What does an assistant usually quote from a page?

Here we move into the third bucket: inference. Let us say plainly that none of these companies has published a list of criteria for choosing a source, so everything in this section is reasoning rather than documentation.

The reasoning is simple. When an assistant builds an answer it needs a sentence it can quote without cutting it up and without adding explanation. A page that answers the question in its first two sentences has such a sentence. A page that answers in the sixth paragraph, after a warm-up, does not. In our experience that single difference does more for being quoted than anything else, and it costs nothing, because it is better for a human reader too.

Three other things help for the same reason: use one term consistently across the page instead of playing with synonyms; link every claim that needs authority to its source, so that if the model wants to check, the path is right there; and lay out the page with proper headings so the relevant part can be lifted out on its own.

Now the point most people leave out, and we say it because it is this site's own policy. Google writes in that same guide that you do not need to write in a specific way to appear in AI features, that there is no requirement to break content into tiny pieces, and that no special schema is needed. So do not sell the four recommendations above as Google's rules. They are an inference about how a language model builds a quotation out of a page, and they are independently right for a human reader. Those two reasons are enough to do them, and there is no third one.

We run this same rule on this Learn section: every lesson is required to answer in its first two sentences, and those two sentences have to make sense alone, without the rest of the page. The first reason is a reader in a hurry; that an assistant can also lift those two sentences is a side effect and not the goal.

Should you build an llms.txt or not?

Short answer: build one if it costs you nothing, but pay nobody for it, and expect no effect on your Google ranking from it.

The document on Google's side is explicit. In that same generative AI guide, in a section it labels mythbusting, it writes that you do not need to create new machine readable files, AI text files, markup or Markdown to appear in Google Search including its generative capabilities, because Google Search itself does not use them. And it immediately adds that it is completely fine if you decide to create and maintain such files, but doing so will neither harm nor help your visibility or rankings in Google Search, as Google Search ignores them.

The other side of the story is real too, and usually left out. llms.txt is a proposal Jeremy Howard published in September 2024, now in its second version at llmstxt.org. That page says thousands of sites publish the file, documentation platforms generate one automatically, Chrome's Lighthouse audits sites for one as part of its agentic browsing checks, and the AI labs themselves publish llms.txt files for their own developer documentation.

So hold both true sentences at once: Google ignores it, and it is at the same time a live convention that other tools do read. We keep one on this site and we are not taking it down, because it costs us one text file that is regenerated automatically. And we do not sell it to anybody, because we have no measured effect to sell.

If somebody is charging you separately to build this file and promising you a result in Google, Google's own documentation has already answered them.

How do you find out whether you are being seen at all?

You have one place of real measurement and one half-measurement, and the rest cannot be measured at all. Let us say honestly which is which.

The real measurement is your own server log. Every request these crawlers make to your site lands there with a name, a time and a response code. That is where you see which of them came, which pages they fetched, and above all what code they got: a page that returns two hundred to your browser can return 403 to a robot, and your server, firewall or CDN can arrange that without telling you. The fast path at the bottom of this page is exactly that check.

The half-measurement is Search Console. Google points in that same guide to the generative AI performance report in Search Console; what that report does and does not show is explained separately in the Search Console AI report.

And the thing nobody can measure: whether your site was quoted inside a private conversation between a user and an assistant. That conversation is available to no tool. This is why Google itself warns, in the mythbusting section, to be wary of third-party tools that promise ranking success or claim to use internal Google metrics, and writes plainly that no third-party tool has access to their internal ranking or AI systems.

There is also a manual method some people call GEO tracking: ask the assistants themselves about your subject and see which sites they point to. It is useful, but it is not measurement; the answer changes each time, it depends on your account history, and the sample size is one. Keep it as a qualitative observation and never make a decision on a chart you built out of it.

The AI crawler access check sheet

Every row here can be checked on your own site in a few minutes, and none of them needs a paid tool.

Crawler accessone domain

Must be true

  • robots.txt carries the crawler tokens by your decision, not by default
  • the server returns 200 to those user agents, not 403
  • the names of these robots actually appear in the last few days of the log
  • an important page shows its text without JavaScript as well

Silent failure

  • the firewall or CDN blocks unknown bots and does not tell you
  • the page sits behind a login or behind a captcha
  • expecting to see Google-Extended in the log, which never comes
  • buying a GEO package that promises placement in answers

Passing this sheet means the door is open. That anybody walks through it is not guaranteed and cannot be.

So what is actually worth doing today?

In the order that they are worth, rather than the order they are sold in GEO packages.

First, make sure your pages are crawlable and indexable at all. A page that is not in the index is not in any AI answer either. If that part is limping, go to technical SEO and put the rest of this lesson aside for now.

Second, and the most common silent failure we meet in practice: make sure your firewall or CDN is not blocking these robots. Many security services are hard on unknown bots by default, and the result is a site with a completely open robots.txt that in practice returns 403 to OAI-SearchBot. Nobody tells you; you simply never get quoted.

Third, set the tokens deliberately rather than by accident. Look at the table in the second section and decide which ones you want and which you do not. If you do not decide, your server default and your security plugin decide for you.

Fourth, answer in the first two sentences and keep one term stable. That is inference and not law, but it is right for the reader too, so it costs you nothing.

And the thing not worth doing: buying a package called GEO that promises placement in AI answers. Nobody outside these companies has access to their systems, and nobody can guarantee a position they cannot see. If you want the four things above done on a real site, that is what our SEO service is; but if you have a technical team, all four are yours to do, and we would rather you did.

One prediction, which may turn out wrong and which we would rather you remembered we made: until one of these companies publishes a report of its citations, everything sold as GEO measurement is a reconstruction of a guess wearing the clothes of a chart. If such a report ever appears, we will rewrite this lesson that day and the verified date at the top of the page will change.

The fast path, with AI

The fast path here is not asking a model how to get seen in AI answers; that question gets you a fabricated answer. It is collecting three real things in a few minutes: your own robots.txt as these robots see it, the response your server gives to those same user agents, and who actually came in the last few days of your log. Then you hand the raw output to a model so it builds an action list where every item is pinned to a line of that output. A cheap fast model is enough; our current pick is in the <a class="text-link" href="/en/ai/">AI section</a>.

  1. Run the first command on your own domain. It has three parts: the crawler tokens in robots.txt, the server response code for four real user agents, and a count of robot names in the log. Copy all three parts in full.
  2. If you have no access to the log path, skip the third part and write that down as a task in itself: without the log you do not have the only real measurement in this subject, and you have to get it from your host.
  3. Hand the output to the model verbatim with the stage two prompt. Summarise nothing; the entire point is that the model works on the real text rather than on your account of it.
  4. Delete any item the model wrote whose evidence line you cannot find in the output. Especially any item that talks about a ranking factor: no such thing is in that output, because no such thing exists.

Copy-ready recipe

== Stage 1: three checks on your own server

D="example.com"; LOG="/var/log/nginx/access.log"
BOTS='gptbot|oai-searchbot|chatgpt-user|claudebot|claude-searchbot|claude-user|perplexitybot|google-extended|applebot|amazonbot|ccbot|meta-externalagent'

echo "== 1. crawler tokens in robots.txt"
curl -s "https://$D/robots.txt" | grep -iE "$BOTS" || echo "no tokens: the default rules"

echo "== 2. what code the server gives these robots"
for UA in "GPTBot/1.2" "OAI-SearchBot/1.4" "ClaudeBot/1.0" "PerplexityBot/1.0"; do
  printf '%-20s ' "$UA"
  curl -s -o /dev/null -w '%{http_code}\n' -A "Mozilla/5.0 (compatible; $UA)" "https://$D/"
done

echo "== 3. who actually came (log)"
grep -oiE "$BOTS" "$LOG" | sort | uniq -c | sort -rn


== Stage 2: turning the output into an action list (cheap fast model)

The output below comes from three checks on one site: the robots.txt tokens,
the server response code for several user agents, and a count of robots in the
log. Build an action list under these rules:

1. every item must carry an evidence line from this output, quoted verbatim
2. do not write any item that has no evidence line
3. prioritise the contradictions: a robot allowed in robots.txt that did not
   get a 200, or a robot that is allowed and does not appear in the log at all
4. write nothing about rankings and nothing about effects on AI answers

---
{stage 1 output}

Before you trust the output: Keep three things for yourself. First, a 200 in stage two only says the door is open; it does not say anybody walked through it. Only part three says that. Second, anybody can forge a user agent string, so the log count is an upper bound rather than an exact number; if an important decision hangs on it, check the requests against that company's published IP list. Third, the model in stage two will most likely try to write something about the effect of these settings on rankings; delete that item, because there is no evidence line for it in the output and no reference for it outside the output either.

AI in this kind of work

Our position here is short: use the models to read output and to see what they say about your subject, and never ask them how to rise in their own answers. An answer to that question always arrives and is always fabricated, because a model has no access to its own source-selection system.

Tools that actually help

  • Claude Suits stage two of the fast path, because it paraphrases less when you ask for verbatim quotes from the output. Iran is on neither of Anthropic's two supported-countries lists.
  • Gemini For seeing which sources an assistant points to on your subject. The answer differs every time, so treat it as an observation rather than a measurement. Our own entry notes that Iran is not on Google's supported-countries list; the payment route is in <a class="text-link" href="/en/ai/buy/">buying AI access</a>.
  • ChatGPT The most important thing you need from this service is not the model but its crawler documentation page, which explains the difference between OAI-SearchBot and GPTBot. We do not have an entry for this service in our AI section yet, and we have not checked access from Iran either.

Where it backfires

The main risk in this subject is fabricated claims about this very subject. Ask any model what the ranking factors in AI answers are and you get a list that is tidy and convincing and has no reference behind it; the question looks like questions that have answers, so the answer comes out looking like an answer. Google itself warns in the mythbusting section of its guide to be wary of third-party tools claiming access to internal metrics, and writes that no third-party tool has access to their internal ranking or AI systems. The second risk is smaller but real: when you ask an assistant what it knows about your brand, the answer depends on your account history and on that session, so do not report the result of asking once as the general state of your brand.

Sources: Google Search Central: optimizing for generative AI features, mythbusting section OpenAI: overview of OpenAI crawlers Anthropic: does Anthropic crawl data from the web, and how can site owners block the crawler

Where this advice stops

Everything this lesson said about what gets quoted is inference and not documentation, because none of these companies has published its criteria for choosing a source. The only certain part is the crawler access layer, and even that is only about opening the door and not about the result: opening it guarantees nobody walks through. The field also moves fast; tokens and policies may differ in a few months, and the verified date at the top of this page is there for exactly that reason. And if your site is not being indexed, none of this is relevant to you yet.

From our own work

The numbers below come from this server's own log on the day this lesson was written, in the window from 00:17 to 08:53 on 5 September 2026 Tehran time, over 68,896 log lines. This site's robots.txt allows sixteen AI crawlers by name, and in that window these came: Applebot with 346 requests, Amazonbot with 239, OAI-SearchBot with 93, meta-externalagent with 84, ChatGPT-User with 46, PerplexityBot with 8 and ClaudeBot with 4. For comparison, Googlebot made 133 requests in the same window and bingbot 49. Two things in that table are useful to you. First, GPTBot is at zero and OAI-SearchBot at 93: those are the two independent tokens from the second section, in practice, on a real site. Second, Google-Extended is at zero too, and that is not a fault; Google's documentation says the token has no separate user agent string, so it never appears in any log. Let us also state what we cannot show you: one measured citation. No tool gives us that, and anybody claiming to have it should first explain where from.

Real follow-up questions

Is GEO different from SEO?

From Google's point of view, no: its official guide says optimizing for generative AI search is optimizing for the search experience and is therefore still SEO. What is genuinely new is the crawler layer: services other than Google have separate tokens and can be configured separately.

If I block GPTBot, do I disappear from ChatGPT?

No, they are independent settings. OpenAI writes that GPTBot is for model training and OAI-SearchBot for being visible in search, and that sites opted out of OAI-SearchBot are not shown in ChatGPT search answers. So if you want to be in search but not in training, block GPTBot only.

Why does Google-Extended never appear in my log?

Because it has no separate user agent string. Google's documentation says crawling is done with the existing strings and that this token acts purely as a control in robots.txt. Its absence from the log is not a fault, and it has nothing to do with your ranking in Search.