AI content and SEO
Google does not penalise a piece of text for having been written with AI; its written policy is about generating many pages that add nothing for the user, no matter who or what produced them. The real risk does not live in the sentences, it lives in the shape of the site: dozens of pages with one template, one date and nothing first-hand.
- Lesson 14 of 15
- Intermediate
- Free, no signup
From assistance to bulk production: the risk spectrum
We placed each mode according to the text of Google's policies. The turning point is the number of pages and the value of each, not the tool that wrote them.
- Help writing one pageoutside the spam policy
- Machine draft, edited and sourcedoutside the spam policy
- Bulk publishing, light editthe borderline
- Bulk generation from one templatescaled content abuse
These positions are our reading of the policy text, not a measured score. Nobody outside Google can produce such a score.
Last checked: Facts and tool names in this lesson are re-checked against their sources on this date.
What has Google actually said about AI-written text?
Google's position has been written on the official Search Central blog since February 2023 and has not changed since. The heading of that page's main section answers the question by itself: rewarding high-quality content, however it is produced.
The sentence to memorise is this one: using automation, including AI, to generate content with the primary purpose of manipulating ranking in search results is a violation of the spam policies. Split it in half. The first half does not condemn the tool. The second half condemns the purpose.
In the FAQ of that same page Google gets blunter. First: using AI does not give content any special gains, it is just content. And second, which we think settles the decision for most businesses: if you see AI as an inexpensive, easy way to game search rankings, then no.
So AI content is not a separate category in Google's vocabulary at all. The category that does exist has a different name, and the next section is about that one. You can read the full position in Google's own guidance about AI-generated content.
Is the line drawn on the tool or on the number of pages?
On the number of pages and the value of each one. In March 2024 Google announced three new spam policies, and one of them lands exactly here: scaled content abuse. The spam policy document defines it as generating many pages whose primary purpose is manipulating search rankings rather than helping users, typically producing large amounts of unoriginal content that provides little to no value, no matter how it is created.
Google explains in that same announcement why the policy was new: the previous one was about automatically generated content, and this one replaced it so that it makes no difference whether the content came from automation, from human effort, or from some combination of the two. The escape route many people were looking for was closed before they got there.
The examples the spam document gives are worth reading, because not one of them is about writing style:
- using generative AI tools to produce many pages without adding value for users
- scraping feeds, search results or other people's content into many pages, including through automated translation or synonymising
- stitching together pieces of several pages without adding value
- creating multiple sites to hide the scaled nature of the content
Read the list once more and notice that no row asks who wrote the text. Every row asks how many, and what was added. That difference is the difference between one page drafted with a model and a two hundred page shelf produced from one template with the keyword swapped; the second is covered by this policy even if a person typed every page.
Put the same thing on one line and you get four modes. Help writing one page sits outside the spam policy, and so does a machine draft, edited and sourced. Bulk publishing, light edit, is the borderline, and the answer there depends on what each page actually added. And bulk generation from one template is the scaled content abuse the policy was written about.
The reverse holds too. If you have a page that tells a real experiment, a measured number or a failed attempt, and the model only wrote the first draft, no clause of this policy touches you. The full definition is in Google's spam policies.
Why does the pattern show even with no detector involved?
This section is the most important part of the lesson and the part least often written in Persian.
Everybody worries about sentences: does this paragraph sound machine written. But what a crawler sees without any special tooling is not a sentence, it is the shape of the site. Picture a site that had twelve pages for eighteen months and then publishes a hundred and sixty new ones in a week. Every one of them has exactly four second-level headings. Every one opens with a definition sentence and closes with a summary paragraph. Every one carries the same internal link block at the end. Not one of them contains a measured number, a client name, a real date or a linked source.
None of that needs a text detector. Publish dates, heading structure, template similarity, the internal link graph and the amount of textual overlap are things any search engine already computes, for other reasons. The pattern is visible in data that has been collected for years.
We cannot see inside Google's systems and nobody outside Google can either; take that sentence seriously, because anyone claiming otherwise is selling a guess. What we can say is this: the written policy names quantity and value rather than style, and quantity and value are exactly the two things that are visible from outside in the pattern above.
The practical consequence is a simple rule. If you want to know whether your publishing carries risk, do not look at the text; look at your list of URLs. If the dates, the lengths and the structures all resemble each other, and no page carries something only you know, the risk is right there, even if you typed every sentence yourself.
What you worry about, and what is actually visible
Above the water sits one paragraph. Below it, five things any crawler already records, none of which need a text detector.
The tone of the paragraph
The only thing everyone worries about, and the only one a rewrite changes.
-
The publish dates of dozens of pages
One week, a hundred and sixty pages, on a site that used to publish one a month.
-
One template for all of them
The same heading count, the same opening sentence, the same closing paragraph.
-
A repeated internal link block
The internal link graph is built anyway; repetition does not hide inside it.
-
Textual overlap between pages
When the subject changes but ninety percent of the sentences do not.
-
The absence of anything first-hand
No number you took yourself, no source, no case where it did not work.
This list does not say what Google weighs; it says these five are visible from outside. The difference between those two sentences matters.
Can you trust an AI text detector?
Not enough to make a decision on, and it is Google DeepMind that wrote that, not us.
On the same page that introduces the SynthID watermark there is a passage about the detection tools in common use: those tools are usually classifiers, classifiers often only perform well on particular tasks, and when the same classifier is applied across different kinds of platform and content its performance is not always reliable or consistent. The consequence, in that text's own words, is that a piece of writing can be mislabelled, for example incorrectly identified as AI-generated.
So the percentage those tools hand you is meaningless in both directions. Rejecting a writer over a score of eighty-two percent is not defensible, and neither is publishing a weak text because the score was four percent. No clause of Google's policies is about the output of these tools either; the policy is about something else, which you read in the last two sections.
One note for the Persian market, stated honestly as a limit of our own knowledge: we have not seen a published evaluation on Persian text from any of these tools. If you find one, read the sample size and the method before trusting it. Deciding on a number that was never evaluated for your language is guessing about a guess.
Our position is plain: if your editorial process depends on getting past a detector, you have built the process around the wrong thing. The thing you have to get past is a reader who came looking for an answer.
What is SynthID and what does it mean for your site?
SynthID is Google DeepMind's watermark: a digital mark embedded directly into AI-generated images, audio, text or video. According to DeepMind's own page the watermark is embedded across Google's generative AI consumer products, is imperceptible to humans, and can be detected by SynthID's own technology.
How it works for text is interesting and worth understanding. A language model builds text one token at a time and gives each candidate a probability score. SynthID nudges those scores at the moment of generation, and that pattern of nudges is the watermark itself. For images and video the mark is added as the content is created, and is designed to survive cropping, filters and lossy compression.
For checking, the official page gives two routes: upload the file to Gemini and ask whether it was created or altered by Google AI, or use the separate SynthID Detector portal, which is currently being tested with journalists and media professionals.
Now the most important part, which the same page states honestly and which is rarely quoted. The text watermark holds up under cropping, changing a few words and mild paraphrasing, but its confidence scores can be greatly reduced when a text is thoroughly rewritten or translated into another language. It is also weaker on short factual answers, because there is less room to adjust probabilities without touching the accuracy of the answer. And DeepMind writes that SynthID is not a silver bullet.
The consequence for you is two sentences. First, a negative check does not prove a human wrote it; it means no watermark was found. Second, none of this is about ranking; it is about whether somebody can later show that an image or a text came out of a Google model. For a site publishing ten generated images a day, that is a reputational fact rather than a ranking factor.
So how do you actually use AI in your publishing?
The rule we run on this site is one sentence: the model gets the drafting and the checking, a person keeps ownership of the facts, the position and the limit.
There are three things a model cannot bring, and those three are what separate one page from two hundred competing ones. A number you measured yourself. A case where this very advice did not work. And the boundary past which the advice stops being true. If a page has none of the three, publishing it only makes the shelf longer.
The order of work we suggest is simple too. First decide what the page carries that exists nowhere else; if you have no answer, stop here. Then take the draft from a model to get there faster. Then attach every factual claim to a source or a measurement, and delete whatever cannot be attached. Last, run the mechanical check, which is in the fast path further down this page.
If you are doing this for a whole site and do not have the capacity, we do this work as content production; but that is your decision and not the conclusion of this lesson.
One last thing experience taught us: for most sites the serious danger is not a Google penalty. The danger is that six months later you own a hundred pages that have no reason to exist, and every change to that site is expensive from then on. The spam policy you can read; that other thing you have to live through to understand.
The sheet we fill in before publishing
The same one written in our own editorial rules. The second column is a reason to reject, not a matter of taste.
Must carry
- a number or observation you took yourself
- a linked source for every claim that needs authority
- a boundary: where this advice stops being true
- a position, even an unwelcome one
Gets it rejected
- a number or quote attached to no source
- a paragraph that would sit in any article on any site
- invisible characters and long dashes from copy and paste
- a flat rhythm: ten paragraphs of the same size in a row
Passing this sheet means the page is publishable, not that it will rank. Ranking is decided by the competition.
The fast path, with AI
The fast path here is not writing faster, it is checking faster. What eats the most time in an editorial process is working out which sentences in a draft carry a factual claim and which of those have nothing behind them; that one job can be handed to a model, because it is mechanical. The three stages below run on any draft, including one you wrote yourself. Stage one works with a cheap fast model, stage two needs a stronger one, and stage three needs no model at all. Our current pick among models is written up in the <a class="text-link" href="/en/ai/">AI section</a>.
- Run stage one with a fast model. Its output is a table: every sentence carrying a factual claim, the kind of claim, and the source present in the text itself. Every row whose source column says none is today's work.
- Every row without a source either gets attached to a real source or gets deleted. There is no third option, and that strictness is the whole value of this path.
- Run stage two on the stronger model and ask it for neither praise nor a rewrite. Ask only what the text is dodging. Judge its answers yourself; some will be wrong, and even those are useful, because they show where a reader gets stuck.
- Run stage three on the final file. This one needs no model and does its job quietly: it finds the invisible characters and long dashes that came in with copy and paste. Silence means clean.
Copy-ready recipe
== Stage 1: claim extraction (cheap fast model)
The text below is an unpublished draft. Pull out every sentence that carries a
factual claim (a number, a percentage, a date, a product or company name, a
quotation, a cause-and-effect claim). Output only a table with three columns:
verbatim sentence | kind of claim | source present in this same text
Rules: do not rewrite any sentence. Do not complete any claim from your own
knowledge. If no source is present in the text, write "none". Add nothing.
---
{draft text}
== Stage 2: what the text is dodging (strong model)
You are a demanding editor in the field of {subject} and this text is about to
be published. Give three things and nothing else:
1. five questions a reader still has after reading this
2. the places where the text makes a general claim but brings no specific case
3. one limit or condition the text does not state, whose absence makes the
text more optimistic than reality
Do not praise it and do not offer a rewrite.
---
{draft text}
== Stage 3: mechanical check (no model)
grep -nP '[\x{200B}-\x{200F}\x{202A}-\x{202E}\x{2060}-\x{2064}\x{FEFF}\x{00A0}\x{00AD}\x{2012}-\x{2015}\x{2212}\x{2018}-\x{201F}\x{2026}\x{2032}\x{2033}]' FILE
Empty output means clean. Every line printed is an invisible character or a
long dash that arrived with copy and paste and has to be replaced with a plain
space or a plain character.
Before you trust the output: Two things stated plainly. First, the point of stage three is not to hide the use of AI and it does not do that; those characters are copy-and-paste rubbish and in Persian they sometimes damage the text as well. A text that passes all three stages and still says nothing new is exactly what the spam policy was written about. Second, the output of stage two is a model's opinion and not a fact; judge it yourself, and never move a new claim the model added in that answer into the text without a source.
AI in this kind of work
In this subject AI is both the tool and the topic, so we will state the position plainly: yes for drafting and for checking, no for deciding what is true. These are the three tools we use ourselves, with what their own pages say about access from Iran.
Tools that actually help
- Claude Suits the claim-extraction stage, because it paraphrases less when you ask for verbatim quotes. Iran is on neither of Anthropic's two supported-countries lists.
- Gemini For a SynthID check this is the only official route: upload the file and ask. Our own entry notes that Iran is not on Google's supported-countries list; the payment route is in <a class="text-link" href="/en/ai/buy/">buying AI access</a>.
- NotebookLM When you have several source documents and want the answers to come only from them, this shape of work reduces the risk of invented claims. You still own the claim, not the tool.
Where it backfires
Put the three specific risks of this subject side by side. First the watermark: the image output of Google's models carries the SynthID mark, that mark is imperceptible to the human eye and detectable by Google's own technology, so a site publishing masses of generated images is identifiable in that respect. The same mechanism runs for text in the Gemini app and web experience. Second the detectors: DeepMind itself writes that classifiers often only perform well on particular tasks and that their performance across platforms and content types is not always reliable, with mislabelling as the result, so put no editorial decision on their number. Third the data: a client's unpublished draft is their text and not yours; whether a conversation is used for model training depends on that service's plan and settings and is written on its own page. None of these three is a Google ranking factor, and we make no such claim.
Sources: Google Search Central: guidance about AI-generated content Google Search Central: spam policies, scaled content abuse Google DeepMind: watermarking AI-generated text and video with SynthID Anthropic: is my data used for model training
Where this advice stops
Passing the spam policy is the floor of the work, not a ranking lever; a page that violates no clause and says nothing new still does not rank. This lesson is only about Google's written policies, too: we cannot see inside the ranking systems and nobody outside Google can, so any sentence claiming that Google detects machine text in some specific way is a guess, whoever it came from. And none of this applies to the other engines; ChatGPT and Perplexity have policies of their own.
From our own work
We run this lesson's position on this very site, and both halves of it can be checked. First, we keep written editorial rules that carry a list of forbidden characters, and a linter runs them over every text before publishing; the strangest row in that list is the zero-width non-joiner, which is correct Persian typography and which we banned on purpose, because it is an invisible character and flawlessly uniform use of it is itself a sign of machine-written text. The decision was not free: we put an ordinary space there instead, and in exchange every text gets read once by a person. Second, this Learn section's release clock: three lessons a day, in three fixed slots at 09:40, 14:50 and 20:10 Tehran time, plus a shift of up to twenty minutes computed from the hash of the lesson's own slug, so no two lessons take the same moment and a hundred and sixty-eight lessons never share one date. The page you are reading is in that queue. This is exactly the pattern described in the third section of this lesson, seen from the side of someone who has to avoid it.
Real follow-up questions
If I write an article with AI, will Google penalise me?
Not for the tool. Google's written policy is about generating many worthless pages, whether they came from automation or from human effort. One article that carries something first-hand and links its sources sits outside that policy.
How accurate are AI text detectors?
Not accurate enough to decide on. Google DeepMind writes on the SynthID page that these classifiers often only perform well on particular tasks, that their performance across different content and platforms is not always consistent, and that a text can be mislabelled. For Persian we have seen no published evaluation at all.
Can I publish AI-generated images on my site?
Yes, but know one fact: the image output of Google's models carries the SynthID watermark, imperceptible to the eye and detectable by Google's own technology. That is not a ranking factor, but it does mean that the generated bulk of a site's images can be demonstrated later.