MONA explainer 11 min read

What Is AI Content Provenance and How SynthID Watermarking Actually Works

Invisible watermark signal embedded across pixels and audio waveforms of AI-generated media

ELI5

AI content provenance is a verified record of a file’s origin and edit history. AI watermarking hides a detectable signal inside the file itself. SynthID does the second part — embedding that pattern into pixels, audio, or token choices.

Upload the same AI-generated photo to Instagram twice — once with its embedded creation record intact, once with that record stripped out by hand. After the platform’s recompression pass, both files come back pixel-different from the original. Only one of them still tells a detector what it is. The other lost its memory somewhere between upload and re-encode, and the reason has nothing to do with the photo itself.

The Pixel Remembers What the Metadata Forgets

That asymmetry isn’t a bug in one platform’s upload pipeline. It’s the structural consequence of two systems solving different problems under the same marketing label. One modifies the content. The other describes it.

What is AI watermarking and content provenance?

Digital Watermarking embeds a detectable signal directly inside generated content — a pattern in pixel values, an inaudible shift in an audio spectrum, a bias in which tokens a language model selects. The signal travels with the bits. Crop the image, recompress the audio, copy half the text into a new document, and the pattern can often still be recovered, because it was never separate from the content to begin with.

Content Provenance solves an adjacent but distinct problem: establishing a verifiable chain of custody for who or what created a piece of media, and what has been done to it since. The leading standard, the Coalition for Content Provenance and Authenticity (C2PA), does this with a manifest — a block of metadata wrapped in a Digital Signature and bound to the file through Cryptographic Hashing, so that any edit invalidates the signature and surfaces as tampering.

Watermarking changes the content; provenance only describes it.

That distinction explains why the two systems fail in opposite directions. A manifest is attached metadata — exactly the kind of field a platform’s re-encoding step routinely discards, the same way EXIF data quietly disappears when you screenshot a photo. A pixel-level watermark is none of that. It’s encoded in the values a re-encoder treats as the actual image, which means destroying the signal means destroying the image too.

The Math Hidden in a Token’s Probability and a Pixel’s Noise

Google DeepMind built SynthID around exactly this kind of pixel-level signal, and the same underlying idea — modify the generation process itself rather than stamp the output afterward — extends across images, audio, video, and text. The mechanism differs by modality. The principle doesn’t.

How does SynthID embed an invisible watermark into AI-generated images and audio?

For images and video, SynthID introduces an imperceptible signal into the pixel values during generation — small enough that human vision doesn’t register it, structured enough that a detection model can recover it after the file has been cropped, recompressed, filtered, or otherwise reprocessed (DeepMind). For audio, the equivalent signal lives in frequencies a listener doesn’t consciously track, embedded directly into the waveform rather than appended as a separate track, which is why it survives format conversion and moderate speed changes.

Not a stamp added after generation. A property of the generation process itself.

That’s the structural reason cropping and recompression don’t erase it the way they erase a manifest: the signal isn’t stored next to the pixels, it’s expressed through them — the same logic that makes Steganography resistant to casual removal in any domain, not only generative AI. The text version of this approach was peer-reviewed in Nature in 2024 (Dathathri et al.), and DeepMind later published a separate technical paper detailing a comparable method for images.

And the Same Idea, Applied to a Token Instead of a Pixel

Text has no pixels to perturb, so SynthID’s text watermark works one layer down, inside the sampling step itself. A logits processor — the component that converts a model’s raw output scores into the probabilities it samples from — nudges those probabilities using a pseudorandom function computed over the vocabulary, looking by default at sequences of five tokens at a time (Google AI Docs). No retraining. No change to the model’s parameters. The bias only shows up in aggregate, across enough tokens, which is exactly what a detector statistically tests for.

The reference implementation is open source, and a production-ready version is now part of Hugging Face’s Transformers library — meaning the technique no longer requires a research lab’s internal tooling; it’s a library import.

A second, unrelated technique works the opposite direction. Perceptual Hashing fingerprints a piece of content’s structure at detection time, which can flag near-duplicates and reposted copies even when no watermark was ever embedded. Useful, but it answers a different question — not “was this generated,” but “have I seen something like this before.”

Diagram comparing how a SynthID pixel-level image watermark and a SynthID text logits-processor watermark each embed and survive editing, next to a C2PA metadata manifest that does not
Two different ways AI content carries an invisible record — one embedded in the content itself, one attached as signed metadata.

What the Signal Predicts, and Where the Chain Breaks

If watermarking and provenance solve different problems, they also fail on different timelines. That has direct consequences for what any given “verified” badge actually tells you.

If a system checks only the embedded signal, expect it to keep working through edits that would defeat a metadata check — a crop, a recompression, a filter pass, a moderate speed change to an audio clip. That’s precisely the resilience SynthID was built for, and the property a research attack is most likely to target first, because it’s the property that actually matters under adversarial conditions.

If a system checks only a C2PA manifest, expect that signal to disappear the moment the file passes through almost any consumer upload pipeline. Major social platforms routinely strip this metadata during their own re-encoding step, which means the badge can vanish before a single human ever tries to remove it.

Neither check, alone, proves where content came from. The practical answer the industry converged on through 2026 is to stop treating the two as alternatives: pair an embedded watermark with a signed provenance manifest, so a missing manifest at least raises a question even when the underlying pixel signal still confirms AI origin.

Rule of thumb: read a watermark hit as evidence of generation, and a provenance manifest as evidence of an unbroken edit history — neither one, alone, is proof of either.

When it breaks: published academic research has demonstrated that an attacker who can systematically manipulate an image — averaging noise patterns across multiple generated samples of it — can collapse a watermark detector’s accuracy far below its claimed reliability. That finding comes from outside DeepMind, not from a flaw the company has disclosed itself, and it means watermark detection should be read as a strong probabilistic signal, not as forensic proof.

Security & compatibility notes:

  • SynthID image-watermark attack: UnMarker (University of Waterloo, presented at IEEE S&P 2025) reports cutting SynthID’s image-detection accuracy from 100% to roughly 21% by averaging noise patterns across multiple samples of the same generated image. The finding comes from the paper itself, not from an official DeepMind disclosure.
  • SynthID text-watermark fragility: Detection accuracy drops sharply after paraphrasing, translation, or substantial copy-paste editing of the underlying text.
  • C2PA manifest stripping: Instagram, X, LinkedIn, TikTok, and Facebook systematically strip C2PA manifests during upload and re-encoding, independent of any user action.
  • C2PA hardware signing: A 2025 vulnerability in one camera maker’s in-device C2PA signing forced revocation of all issued certificates; the service had not been restored as of early 2026.

The Race to Make Provenance the Default

The mechanism is half the story. The other half is who actually turns it on.

By 2026, watermarking stopped being an optional feature bolted onto a model and became default behavior across the Generative Media APIs that generate the content in the first place — Gemini, Imagen, Lyria, and Veo all embed SynthID automatically, with no separate setting to enable (DeepMind). SynthID has labeled more than 100 billion images and videos, plus the audio equivalent of 60,000 years, since its 2023 launch — a Google-reported figure that, per InfoQ, hasn’t been independently audited. OpenAI joined the C2PA steering committee in May 2026 and now embeds SynthID alongside C2PA metadata in its own outputs; ElevenLabs, Kakao, and Nvidia have adopted it too (InfoQ).

The provenance side of the standard grew the same way, through a different mechanism — coalition membership rather than API defaults. As of December 2025, the published C2PA technical specification sat at version 2.3, with a 2.4 revision already out (C2PA spec) — a number worth checking again before quoting it as current, since this category moves fast.

Dedicated vendors fill the space around the two majors. Truepic signs media with verified time, date, device, and location at the moment of capture or generation, anchoring the provenance chain at the camera or generator instead of at upload. Steg.AI sells a comparable invisible-watermarking capability as an enterprise API, aimed less at labeling AI content than at tracing how a specific leaked file made its way out of a company.

None of this stays confined to images or audio in isolation. Across the broader Generative Media Pipelines stack, the direction is the same: watermarking and provenance metadata are becoming something a pipeline emits by default, not a feature a developer has to remember to request.

The Data Says

A watermark proves origin at the bit level; a provenance manifest proves custody at the metadata level — treat either one alone as partial evidence, not verification. SynthID’s resilience against everyday edits is real and independently attackable in a research setting; C2PA’s manifest is independently strong and trivially erased by the upload pipelines most people actually use. The two systems were built to compensate for each other’s blind spot, which is the only way either one is currently worth trusting.

AI-assisted content, human-reviewed. Images AI-generated. Editorial Standards · Our Editors