AI image watermarks and SynthID
An AI watermark is a mark the model vendor puts inside the output so it can later be told that a machine made the file, and Google own documentation says every image the Gemini API generates carries one. It is invisible to the eye, but anyone holding the file can look for it.
- Lesson 12 of 12
- Intermediate
- Free, no signup
What you see in the image and what is in it
The image you publish
The pixels. The only layer visible when you open the file, and the only thing a visitor sees.
-
The in-pixel watermark
A pattern the model places in the image itself while generating. Part of the image, not something attached to it. SynthID lives here.
-
The list of actions
Created, converted, edited, plus the name of the software that performed each action.
-
The digital source code
A code from the IPTC vocabulary. For a fully machine-made image it is trainedAlgorithmicMedia.
-
The chain of earlier claims
A reference to the credential of the previous version, so converting the file does not cut its history.
-
Signature and certificate chain
Who made the claim and which authority vouches for the signature. Without it the other layers are just text.
None of these layers is visible by looking at the image, and the bottom four are all metadata: one automatic resize can remove all four completely without raising any error.
Last checked: Facts and tool names in this lesson are re-checked against their sources on this date.
Two different things both called a watermark
An AI watermark is not that logo in the corner of an image. Two entirely different mechanisms sit under one word, and the difference is exactly what changes your decision.
The first lives inside the pixels. While generating, the model places a pattern in the image itself that the human eye does not see and that does not change quality. Because it is part of the image, copying or re-saving does not remove it. Google SynthID is this kind.
The second sits beside the image rather than in it: a signed data package attached to the file, saying what made this file, what edits it went through, and who signed that claim. The standard is called C2PA and what it produces is called Content Credentials. This one is removable like any other metadata, and in practice it is removed far more easily than people expect; section five shows exactly that on a real file.
So an image can carry no credential and still carry a watermark, or carry a credential and no watermark at all. Treat the two as one thing and you will misread every result a checking tool gives you.
What SynthID does and where it is
On its own page, Google DeepMind writes that SynthID embeds digital watermarks directly into AI-generated images, audio, text or video, that those watermarks are embedded across Google generative AI consumer products, that they are imperceptible to humans, and that they can be detected by SynthID own technology.
For images and video the same page says the watermark is added the moment the content is created and does not change quality. For audio it is embedded in anything generated through the Lyria music model or the NotebookLM podcast feature. For text, Google says it extended SynthID to output from the Gemini app and web experience, and that it works by nudging the probability score of the next word so a pattern remains in the word choices themselves.
But the line a designer needs is written somewhere else, and more bluntly. The Gemini API image generation documentation carries one sentence that settles the matter: all generated images include a SynthID watermark. It is not an optional feature you can switch off; Nano Banana 2 and Nano Banana Pro output comes out with it.
Two things Google has not said, and neither do we: no number has been published for what share of images stays detectable after any given edit, and nowhere is it claimed that models from other vendors carry this mark. What is above is Google own claim about Google own products, and we quote it as exactly that.
What Content Credentials are and who attaches them
C2PA is an open standard published by a coalition of the same name, and what it produces is called Content Credentials. The coalition own site gives it a good analogy: like a nutrition label on a package, it puts the history of a piece of content within anyone reach.
Inside that package sit a few specific things. A list of actions performed on the file, such as created and converted. The name of the software that did it. A code from the IPTC vocabulary saying what the digital source of this file was, and the code used for a fully machine-made image is trainedAlgorithmicMedia. A signature and a certificate chain saying who made the claim. And a reference to earlier claims, so that converting a file does not cut its history off.
The names standing behind the standard are not small ones. On the day this lesson was written, the coalition membership page listed eleven steering committee members: Adobe, Amazon, the BBC, Google, Meta, Microsoft, OpenAI, Publicis Groupe, Sony, TikTok and Truepic. Model vendors, social networks, broadcasters and camera makers all at once.
One thing should be said plainly because it gets confused everywhere: sitting on the coalition is not the same as attaching a credential to your output. A company being on the steering committee does not mean all of its products write Content Credentials. The only thing you can say with confidence is what a vendor has documented, or what you can see by opening a real file; and in section five we do exactly that on a file from this site.
In-pixel mark versus signed credential
In-pixel watermark, such as SynthID
- Part of the image itself, not attached to it
- Its vendor says it is designed to survive cropping and compression
- Reading it needs that same vendor tooling
- Says nothing about later edits
Signed credential, meaning C2PA
- A data package attached to the file
- Editing, converting or sharing may remove it
- Readable with a public web tool and no account
- Tells you the whole edit chain and who signed it
Neither replaces the other, and one image can carry one, both or neither. Both columns are quoted from vendor documentation; we have measured neither.
Who can actually read these marks
This is where most write-ups get it wrong, because the answer is not the same for the two mechanisms.
For SynthID, the route open to everyone is the one Google itself describes: upload the file into a Gemini chat and ask whether it was created or edited by Google AI. Gemini looks for a SynthID watermark and tells you. Alongside it Google has a separate portal called SynthID Detector, but that same page says it is currently collaborating with journalists and media professionals to test the portal and collect feedback, and offers an early tester waitlist. A general, open detector for everyone is not a thing that exists today.
For Content Credentials the situation is more open. The Content Authenticity Initiative runs a public web tool where anyone can drop a file and see its credential. OpenAI has also published a content provenance API that checks both at once, Content Credentials and SynthID, with a web tool alongside it.
And now the part that matters more than any tool, written by OpenAI in its own documentation: a not detected result means the tool found no supported signal, not that the content is human-made. The same page explains that metadata may have been stripped, a watermark may have been degraded, the file may have come from a legacy model, and it states plainly that the tool does not currently detect content generated by another company AI model. A clean check clears nothing.
The reverse holds too: finding a signal is evidence of that signal, not a complete history of the file. The same documentation says to check the signature issuer before attributing an image to a particular provider. These tools collect evidence; they do not deliver verdicts.
What survives editing and uploading
Put the vendor claims on the table as claims first. Google writes that the image and video watermark is designed to stand up to modifications like cropping, adding filters, changing frame rates and lossy compression. OpenAI documentation has a more precise sentence, and it explains the difference between the two layers completely: C2PA metadata provides more context about a file origin, but editing, converting or sharing a file can remove its metadata; a SynthID watermark is part of the image or audio itself and may survive some transformations.
Now look at the same thing on a real file, because the difference between can be removed and is removed only shows up when you open the bytes.
On the day this lesson was written we scanned all 3,176 image files in the uploads folder of this site. Fifty-four carried a C2PA credential. Fifty-three of them were files uploaded by users of the site own freelance marketplace, and fifty of that set carry exactly the code trainedAlgorithmicMedia, meaning the file itself says it is entirely machine-made.
The fifty-fourth is ours. The featured image of the website SEO page, added to the media library in June 2024, carries a credential saying it was made with DALL-E, its digital source code is that same trainedAlgorithmicMedia, it was converted once after creation, and its signature leads to a certificate chain in the name of Truepic. We did not put that there; it had been sitting there and nobody had looked.
Here is the interesting part. WordPress generated twenty-seven other sizes from that file and not one of those twenty-seven carries a single byte of that credential. No credential, no metadata packet. One automatic resize wiped out the very thing meant to provide transparency, and nobody saw an error.
The consequence is the same for any WordPress site, and it cuts both ways. One way: do not count on a credential to disclose for you, because on most publishing paths it never reaches a visitor at all. The other way: metadata being wiped does not mean the trace is wiped, because the in-pixel layer is a different thing and we have no tool to read it.
And one exception that happens to be interesting here: the featured image of that page is published as its social sharing image, and it is the original that gets served, not the resized versions. So every time that page is shared in Telegram, WhatsApp or a social network, the file downloaded at the other end is the one that carries the credential.
The path a credential takes through WordPress
-
1
Generation
The service attaches the credential: the action, the software, the digital source code and the signature.
-
2
Upload
The original lands intact on the server, its credential still complete inside it.
-
3
Automatic resize
WordPress builds several smaller versions. None of them carries the credential, and no error is raised.
-
4
The page picks a version
The browser usually takes one of those resized versions rather than the original.
-
5
The visitor
What gets downloaded has no credential. The transparency the service added never reached anyone.
This chain is only about the credential. We have no tool for reading an in-pixel watermark, so we make no claim about whether SynthID survives a resize.
What this means for your site in search
First, what we do not say, because it is said everywhere without a source. Google has nowhere stated that it reads a SynthID watermark or a C2PA credential as a ranking signal, we have no evidence that it does, and we print no percentage about detection. Anyone who hands you a firm number should be asked where it came from.
What can be said is narrower and checkable: this information is inside your files, and anyone holding the file can read it. A site publishing a hundred images that all carry the code trainedAlgorithmicMedia is classifiable from the files alone, without anyone even looking at the page text. That is not a claim about ranking; it is a fact about identifiability.
Now the part that is actually written in Google policies. The Google Search spam policies have a section called scaled content abuse, and its definition is this: many pages generated for the primary purpose of manipulating search rankings and not helping users, typically focused on creating large amounts of unoriginal content that provides little to no value to users, no matter how it is created. One of its examples is explicit: using generative AI tools to generate many pages without adding value for users.
Google own guidance on AI-generated content says the same thing from another angle: using automation, including AI, to generate content with the primary purpose of manipulating ranking in search results is a violation of the spam policies; and it immediately adds that not all use of automation is spam.
So the thread connecting the two ends is mechanism, not punishment. A page whose content is unoriginal and mass-produced has a problem because it is unoriginal and mass-produced. The watermarked images on that page create no new signal; they only make what you did readable from the files themselves, for Google and for everyone else. If the page content genuinely has value, that readability does you no harm. And if it does not, the images are not your problem.
To see where the same argument lands for text, the lesson on AI content and SEO opens it from the text side.
Our position, and what this lesson will not do
In client work, say that AI was used. One sentence in the contract or in the handover is enough: which images came from a model and which are photographs or hand work. That sentence costs nothing and buys the only thing that does not come back once it is gone.
Put machine-made images where originality is not the subject: drafts, mood boards, backgrounds, explanatory shapes, internal decks. Spend where originality is itself the product: product shots, team photos, brand identity, portfolio pieces. The featured image on our own SEO page, the one above, falls squarely in the first group; we have kept it and we have written here what it is, which is what we would expect of anyone else.
If you made an image with AI and it is going to make a real claim, replace the image rather than the credential. Those two acts do not have the same result: one solves the problem, the other only erases the evidence of it.
And for that reason this lesson does not write down how to remove a watermark. The layer exists so a reader can tell what they are looking at, and teaching its removal is teaching deception. There is a technical point too, rarely made: stripping metadata and stripping an in-pixel mark are two different jobs, and whoever strips the metadata usually believes they stripped both.
The fast path, with AI
The fast path here is not dropping images one by one into a web tool. What a professional does is <strong>scan the whole media library once</strong>, because a C2PA credential is a data package with a specific textual signature and finding it needs no service at all. The command below is the one we ran on this site, and on 3,176 files it took under a tenth of a second. For writing the disclosure sentence at the end, a fast model is enough; our current pick lives in <a class="text-link" href="/en/ai/">the AI section</a>.
- Run the first command in your site uploads folder. The output is the list of files that carry a credential. Compare the size of that list with your total number of image files, not with your expectation.
- For every file found, run the second command and read its credential. Look at three things: the name of the generating software, the digital source code, and which actions are recorded on the file.
- Put every file in one of two groups: decorative, where keeping it is fine, and claim-making, an image a customer makes a decision from. Replace the second group; do not erase its credential.
- Write one disclosure sentence and put it in a fixed place on the site or in the client handover. It should say which class of images came from a model, not that AI is used in general.
- Put the same scan into your upload routine. If your site accepts user uploads this is not a one-off job: fifty-three of our fifty-four files came from users.
Copy-ready recipe
# 1. Search the whole media library for a C2PA credential.
# Run this inside your own site uploads folder.
grep -rlaiE 'c2pa|jumbf' \
--include='*.jpg' --include='*.jpeg' \
--include='*.png' --include='*.webp' .
# For comparison, the total number of image files:
find . \( -iname '*.jpg' -o -iname '*.jpeg' \
-o -iname '*.png' -o -iname '*.webp' \) | wc -l
# 2. Read the credential of one of the files it found.
strings -a {path to the file} \
| grep -iE 'softwareAgent|digitalsourcetype|claim_generator'
# The three things to look for in the output:
# softwareAgent the name of the tool that made the image
# digitalsourcetype the IPTC code; trainedAlgorithmicMedia means fully machine-made
# claim_generator the name of the service that signed the claim
Before you trust the output: This scan finds only the C2PA credential, not the in-pixel watermark. So an empty result proves nothing: a file with no credential at all can still carry a SynthID watermark, and this command will never tell you. Two more things are a person job. Do not read the count alone; actually open the credential, because the difference between a fully machine-made image and a photograph that took an AI edit is written in that digital source code. And the decision whether an image stays or gets replaced is yours, not the output of a command.
AI in this kind of work
This lesson is about what AI puts inside its own output, so the tools here fall into two groups: those that place a mark and those that read one. Our position in one sentence: <strong>none of these tools tells you an image is human-made; they tell you a mark was found or was not, and those are two different statements.</strong>
Tools that actually help
- Gemini The only SynthID check Google itself describes for everyone: upload the file into a chat and ask whether it was created or edited by Google AI. The answer is about Google watermark, not about any other model. Google own page says the Gemini web app works in more than two hundred and thirty countries and territories, and Iran is not on that list.
- OpenAI Content Provenance API The only tool whose documentation says it checks both layers at once: the C2PA credential and the SynthID watermark on an image. For us the most important thing about it is not the capability but the honesty of its docs: it says plainly that a not detected result means no signal was found, not that the content is human-made, and that it does not detect content generated by another company model. It has a web tool too, which we could not open; the main OpenAI domain returns 403 to our server.
- Content Credentials Verify A public web tool for reading the C2PA credential of a file. It needs no account and does exactly what the command line in the fast path does, only with an interface and more readable output. For a single file this is simpler; for three thousand files, use the command.
- Nano Banana 2 It is here because it is the other side of the story: the Gemini API documentation says all generated images include a SynthID watermark, so this model and Nano Banana Pro produce marked output. That is not a defect and for most work it does not matter at all; you just need to know what the file you hand over carries with it.
Where it backfires
The main risk here is not the watermark; it is people treating a checking tool as a verdict.
It goes wrong in both directions. Someone who gets a nothing found result and concludes the image is human-made has said something that tool own documentation refuses: the metadata may have been stripped, the watermark may have been degraded, or the generating model may belong to another company that the tool does not detect at all. And someone who finds a credential and attributes the image to a vendor from it has done exactly what the documentation says not to do: check the signature issuer first.
The more practical risk for an agency is this: leaning on a credential as your disclosure. Section five of this lesson shows one automatic WordPress resize removing the credential completely; so an image you assume introduces itself is in practice telling a visitor nothing. Disclosure is a sentence in your text, not a job for metadata.
And one thing this lesson deliberately does not give: a way to remove a watermark. The reason is in the last section.
Sources: Google DeepMind: SynthID Google: image generation in the Gemini API OpenAI: content provenance guide C2PA: the coalition and Content Credentials C2PA: membership and steering committee IPTC: digital source type, trainedAlgorithmicMedia Google Search: spam policies, scaled content abuse Google Search: guidance about AI-generated content Google: where the Gemini web app is available
Where this advice stops
The boundary of this lesson has to be stated plainly, because its subject is exactly where unbacked claims settle in easily. We read the bytes of the files and we did not run a C2PA cryptographic validator; the tool is not installed on this server. So what we said about that file is the content of its credential, not a verdict on whether the signature is sound. On SynthID we are more limited still: we have no way to detect it at all, so every sentence about SynthID in this lesson is quoted from Google rather than measured by us, including the claim that it survives cropping and compression. Three things are deliberately absent. Any detection rate, because we have no source we would defend. Any claim that Google treats these marks as a ranking signal, because we have no evidence for it and Google has said no such thing. And anything from OpenAI own pages on its main domain, because they return 403 to our server; everything quoted comes from its developer documentation, which does open. One last caveat about the numbers here: our scan is one site on one day, not a survey; do not read fifty-four out of 3,176 as a statistic about the web.
From our own work
We wrote this lesson by looking at our own files, and the first thing that turned up was ours. The featured image of the website SEO page, one of our main service pages, has been in the media library since June 2024 and carries a complete C2PA credential: made with DALL-E, digital source code trainedAlgorithmicMedia, converted once, and a signature leading to a certificate chain in the name of Truepic. It is 182,418 bytes, and that same file is published as the social sharing image of that page, so every time the page is shared in a messenger the credential travels with it. Nobody hid this; nobody had looked. Two more numbers from the same scan across all 3,176 image files on the site: fifty-four carried a credential and fifty-three of those were uploads from users of our own freelance marketplace, which means this information enters the site through routes the content team does not control at all. And of the twenty-seven versions WordPress generated from that same file, not one carries a single byte of the credential.
Real follow-up questions
How do I tell whether an image was made by AI?
It takes two separate checks. For the credential, drop the file into a content authenticity web tool or run the command from the fast path. For the Google watermark, upload the file into a Gemini chat and ask. And in both cases, finding nothing proves nothing.
Does Google penalise sites that use AI images?
Google has nowhere said it reads the watermark, and neither do we. What its own policies do describe is unoriginal, mass-produced content, no matter how it was created. The issue is the quality and originality of the page, not the format of an image file.
If I strip the image metadata, is the problem solved?
No, and in practice it has usually already happened by itself: one resize in WordPress removes the credential without you doing anything. But it solves neither problem: the in-pixel mark is a different layer, and if the image makes a real claim, the metadata was never the problem.