Machine translation leaves a structural fingerprint on text, and it can be measured from the outside. We went looking for it, and the first place we found it was our own site: all 32 published articles on rgb.ir carry exactly the same number of H2 headings in every one of their four languages.
Thirty two out of thirty two. Not mostly. All of them.

What we counted
This site is served in four languages, with Persian as the source and English, Arabic and Turkish sharing the same URL. For every article, in every language, we counted paragraphs, H2 headings, H3 headings, list items, links, words and sentences.
| Measure | Identical in all four languages |
|---|---|
| H2 count | 32 of 32 |
| Link count | 31 of 32 |
| List item count | 31 of 32 |
| Paragraph count | 30 of 32 |
| All five at once | 28 of 32 (88%) |
| Sentence count | 8 of 32 (25%) |
The median spread in sentence count across the four languages was one sentence. One.
Ask yourself the obvious question. If four professional writers each composed the same argument from scratch in their own language, how often would they land on identical heading counts? Not thirty two times out of thirty two.
The word ratios say it even louder
Average word count against the Persian was tight and repeatable too: English 104.5 percent, Arabic 89.3 percent, Turkish 88.4 percent.
Those ratios are not properties of the languages. They are the ratios you get when you convert sentence by sentence from a shared source. Independently written Turkish sometimes runs longer than the Persian and sometimes half its length. Anything that sits at 88 percent every time is tracking the original.
The caveat without which this analysis is incomplete
Identical structure does not prove a machine did it. A disciplined human translator produces exactly the same signature, because fidelity to the source is the job.
So what these numbers demonstrate is not "machine". It is conversion: the second text was built from the first text rather than from the subject. For a device manual, conversion is precisely the right method. For a page that is supposed to rank in Turkey, it is not.
Why any of this affects rankings
Google has not banned machine translation as such. What the spam policies name is translated text published without human review. The distinction is the review, not the tool.
The practical problem is more immediate than the policy. Somebody in Istanbul searching for a service types a phrase that is not the word-for-word rendering of the Persian one. A page converted from Persian does not contain the words they typed, because those words were never part of any decision. The text is not wrong. It just answers a question nobody asked.
That is the thing we accepted about our own site. Our 32 articles are technically flawless in three other languages, and in those three languages nobody had ever measured demand.
How to run this check on your own site
You do not need to buy a tool. Open one page in two languages and count the H2 headings. If they match, count the paragraphs. If those match too, you have your answer.
Then a second and harsher test: look in the English version for an idiom that only English speakers use and that has no direct equivalent in the source language. If you cannot find one, that text was not written.
So what should you do with machine translation
The position we have taken, and pay for: machines to understand, humans to write. Machine translation is excellent and cheap for working out what a competitor's page says. For a page meant to sell, throw it away and start from the subject instead.
And the boundary, because every recommendation has one. If your content is technical documentation, terms and conditions, or product specifications, conversion is correct and rewriting is a waste of money. Sameness there is a feature rather than a defect.
We are rewriting the old articles one at a time rather than retranslating them. It is slow, and the numbers above are the reason.
Is a multilingual site worth it at all, then?
Yes, if each language gets its own content. The technical structure matters too, and Google's guide to localised versions is the right starting point.
Will machine translation get me penalised?
Not by itself. Low quality unreviewed text published at volume, yes. The tool is not the deciding factor.
If you want your language versions written rather than converted, translating and localising site content is what we do and how we do it. If the source-language content still has gaps, written content production comes first, and if the real problem is visibility rather than language, an SEO review of the site is a better place to start. For a site built multilingual from the beginning, this decision belongs to the architecture of the site itself rather than to something bolted on later.
Comments & Questions
Have a question about this article? Ask, we'll answer.
No comments yet; be the first.