(field notes · AI provenance · )

The Race to
Watermark
Reality

OpenAI borrowed Google's secret weapon to detect AI images. Here's exactly how both technologies work — and where they still break.

  • Deep dive
  • ~12 min read
  • 3M images / week via ChatGPT

Three million images generated through ChatGPT in a single week, and almost none of them verifiable once they leave the app.

This is a technical breakdown of the two systems now trying to close that gap: C2PA, a cryptographically signed manifest attached to the file, and SynthID, a neural watermark embedded in the pixel values themselves. How each one is built, what each one survives, where each one fails, and the hole neither of them covers.

01The problem

3 million images. Almost none verifiable.

That's how many images were generated through ChatGPT in a single week — the number that landed May 19, 2026 and put OpenAI on notice fast enough to ship a major announcement.

The problem: almost none of them can be reliably verified as AI-generated once they leave ChatGPT and start travelling across the internet.

They get screenshotted. Reposted. Compressed. Cropped. Saved as different formats. By the time a fake headshot of a CEO or a fabricated protest photo lands on your timeline, the trail is cold.

Why EXIF was never enough

The old solution was EXIF metadata — cameras bake origin information into image files. Problem: any image editor can strip or fake EXIF data in three seconds. It was never designed for adversarial conditions. The internet needed something fundamentally different.

OpenAI's response: adopt Google DeepMind's SynthID watermarking (the exact same system Google already built into Gemini) and pair it with an open standard called C2PA. Then launch openai.com/verify so anyone can check.

But these two technologies work completely differently, fail in completely different ways — and one is orders of magnitude harder to break than the other. Let's go deep.

02Two layers

The passport vs. the tattoo

The dual-layer system OpenAI (and Google) uses combines two fundamentally different approaches. Understanding why they're different is the whole point.

Layer 01 · The passport

C2PA

A cryptographically signed manifest attached to the image file. Founded 2021 by Adobe, Microsoft, BBC, Intel, ARM. It contains who created it, when, what tools, what edits — verifiable against the creator's public certificate.

image file ATTACHED C2PA MANIFEST CLAIM + EDITS SIGNATURE strippable in one command

What it records

  • Claim generator (ChatGPT ImageGen, Nikon Z6 III…)
  • Timestamp + full edit audit trail
  • Cryptographic signature (same tech as HTTPS)
  • Works for AI-made AND camera-made images

Fatal flaws

  • Screenshot destroys it — new file, no metadata
  • Every major social platform strips it on upload
  • Format conversion breaks the chain
  • Anyone can strip it with one ExifTool command
Layer 02 · The tattoo

SynthID

A neural watermark embedded directly into pixel values at generation time. Built by Google DeepMind. Invisible to human eyes. It uses frequency-domain steganography — the pattern lives inside the image content, not attached to it.

INSIDE THE PIXELS ±1–2 RGB per pixel · cryptographic key pattern

What makes it durable

  • Pattern is in the pixels — screenshots capture pixels
  • Survives social media JPEG recompression (~80% quality)
  • Survives resizing, format conversion
  • Cannot be stripped without modifying image quality

Where it breaks

  • Adversarial regeneration attacks (academic, not trivial)
  • Extreme manipulation (crop 90% + aggressive recompress)
  • Only works for images from participating generators
  • Midjourney, local Stable Diffusion — no watermark at all
C2PA · THE PASSPORT SIG the manifest sits beside the pixels strip the metadata and it is gone SYNTHID · THE TATTOO the pattern is the pixels it survives what strips metadata two layers, two completely different failure modes
03The journey

Follow an image across the internet

Step through each stage to see exactly what happens to both signals as an AI-generated image travels from ChatGPT through the real internet.

Step 1 · the source

ChatGPT output

Fresh from the model. Full provenance intact. One file, two complete protection layers.

C2PA · valid — signed by OpenAI · 100%

Manifest signed by OpenAI. Contains model ID, timestamp, generation parameters. Certificate verifiable against OpenAI's public certificate authority.

SynthID · strong — 94% confidence

Watermark embedded across full frequency spectrum at generation time by SynthID encoder network.

Step 2 · the container survives

Downloaded

File download preserves the container intact. Both signals travel with the file.

C2PA · valid — intact · 100%

JPEG container unchanged. C2PA manifest preserved in XMP metadata block — standard file download doesn't touch it.

SynthID · strong — 94% confidence

Pixel data unchanged during download. Watermark at full strength.

Step 3 · the platform pipeline

Instagram

Instagram reprocesses all uploads through their CDN pipeline. C2PA is gone. SynthID survives.

C2PA · stripped — gone · 0%

Instagram removes all metadata including XMP, EXIF, and IPTC on upload. Manifest permanently gone. This is true of Twitter/X, Facebook, TikTok, and LinkedIn too.

SynthID · degraded — 71%

JPEG recompression at ~80% quality degraded signal. Distributed frequency pattern still detectable at 71% confidence — above the 40% threshold.

Step 4 · a brand new file

Screenshot

Screenshot creates a brand new file from your display buffer. C2PA was never in this file.

C2PA · never present · 0%

Screenshot captures display buffer pixels into a new PNG or JPEG. No original file metadata — this is a new file, not a copy. C2PA was never attached to it.

SynthID · degraded — 68%

Pixel data rendered to screen by GPU, then recaptured by the screenshot tool. SynthID pattern survived the display render pipeline at 68% confidence.

Step 5 · past the threshold

Heavy crop

Aggressive manipulation pushes SynthID below reliable detection threshold.

C2PA · never present · 0%

C2PA was gone since step 3. Still not present — and now the file has been manipulated further.

SynthID · below threshold — 28%

Crop removed 80% of watermark-bearing pixels. 5× recompression further degraded remaining signal to 28% — below the ~40% detection threshold. Detector returns “no watermark found.”

C2PA SYNTHID 01 02 03 04 05 100% 100% 0% 0% 0% 94% 94% 71% 68% 28% ChatGPT output downloaded Instagram screenshot heavy crop metadata dies at the first upload · the pixel watermark carries the image three stages further
04Technical mechanism

Inside the frequency domain

Most explainers skip the actual mechanism. Here's what SynthID's encoder does at a technical level — and why it survives what kills metadata-based approaches.

The key idea: steganography

Steganography — hiding a message inside another message — is the key idea. Medieval spies shaved a messenger's head, tattooed a message on the scalp, waited for hair to regrow, then sent them through enemy lines. SynthID does the same thing with pixels instead of scalps.

LOW FREQUENCY ← SYNTHID ZONE (HIGH FREQ) Frequency domain representation: SynthID targets the high-frequency bands (right) that human vision is least sensitive to

After an image is generated by the diffusion model, SynthID's encoder network takes the finished image and applies learned perturbations to pixel values. These aren't random noise — they're structured according to a cryptographic key pattern, distributed across the entire image in the frequency domain.

Human vision is most sensitive to low-frequency changes — broad colours, overall composition. SynthID targets the high-frequency zones — fine textures, sharp edges — where modifications of ±1–2 RGB units are statistically detectable but perceptually invisible.

Pixel-level demo

The three panels show: original image, SynthID-watermarked (±1.5 RGB), and amplified 25× (what the neural detector reads). Three modes of the same ten-by-ten patch, laid out side by side so you can compare them at once: normal view, apply watermark, amplify 25× (detector view).

ORIGINAL no watermark baseline image WITH SYNTHID (±1.5 RGB) looks identical to human eyes at normal zoom AMPLIFIED 25× — DETECTOR neural detector reads this accent cells = correlation

The left panel is the baseline: no watermark. The middle panel carries the watermark and looks identical to human eyes at normal zoom — that is the entire design goal. The right panel is the same watermarked data amplified 25× — bright spots = high correlation with the key. The perturbations were never random, so amplifying them separates the structure out cleanly — a structure the detector already knows how to look for.

SynthID confidence across real-world transformations

Each transformation costs the signal something. Here is what survives what.

Fresh ChatGPT output 94% Gemini Omni image 91% Gemini Nano image 89% Instagram JPEG recompression 71% Screenshot (display re-render) 68% Aggressive crop + 5× compress 28% Adversarial regeneration attack 12% 40% detection threshold below the dashed line the detector returns “no watermark found”

The 40% dashed threshold — signals below it return “no watermark found.”

05The verifier

Run the verification yourself

This is what openai.com/verify returns for different scenarios — including images from Gemini Nano and Gemini Omni verified through Google's own tool.

ScenarioC2PA manifestSynthID watermarkVerdict
Fresh ChatGPT downloadValidDetected — 94%AI-generated (OpenAI confirmed)
Instagram round-tripStrippedDetected — 71%Likely OpenAI-origin
ScreenshotNever presentDetected — 68%Likely OpenAI-origin
Gemini Nano imageValid — Google LLCDetected — 89%AI-generated (Google — not OpenAI)
Gemini Omni imageValid — Google LLCDetected — 91%AI-generated (Google Gemini)
Midjourney imageNot foundNot detectedNo signals found — inconclusive
Heavily manipulatedGoneBelow threshold — 28%Inconclusive — cannot verify

Open any row below for the full verification note the tool returns.

Verdict — AI-generated (OpenAI confirmed). C2PA manifest: valid. SynthID watermark: detected at 94%.

Both C2PA manifest and SynthID watermark detected. Certificate issuer: OpenAI LLC. Image is confirmed AI-generated.

Verdict — likely OpenAI-origin. C2PA manifest: stripped. SynthID watermark: detected at 71%.

C2PA manifest stripped by Instagram CDN. SynthID watermark still detected at 71% confidence — above detection threshold. Likely OpenAI-origin despite missing metadata.

Verdict — likely OpenAI-origin. C2PA manifest: never present. SynthID watermark: detected at 68%.

No C2PA metadata — screenshots create new files. SynthID watermark detected at 68% confidence. Dual-layer system working as designed: metadata fails, pixel watermark catches it.

Verdict — AI-generated (Google, not OpenAI). C2PA manifest: valid — Google LLC. SynthID watermark: detected at 89%.

C2PA manifest signed by Google LLC (not OpenAI). SynthID watermark detected at 89% confidence — encoded by Google DeepMind. Works across Gemini Nano, Flash, Pro, and Omni. Verify at gemini.google.com or via Google Search.

Verdict — AI-generated (Google Gemini). C2PA manifest: valid — Google LLC. SynthID watermark: detected at 91%.

Gemini Omni's higher-resolution outputs carry SynthID at 91% confidence. Full C2PA audit trail with model version, generation timestamp. Google's verification is native in Gemini chat — just upload and ask.

Verdict — no signals found, inconclusive. C2PA manifest: not found. SynthID watermark: not detected.

No C2PA metadata. No SynthID watermark. Midjourney, local Stable Diffusion, ComfyUI, Flux — none of these use SynthID. This is the biggest gap in the system: the most dangerous AI-generated content comes from tools that are completely outside it.

Verdict — inconclusive, cannot verify. C2PA manifest: gone. SynthID watermark: below threshold at 28%.

SynthID signal degraded to 28% — below reliable detection threshold. Cannot confirm or deny AI origin. Bad actors know this: aggressive manipulation is the primary attack vector against pixel-level watermarking.

06What nobody covered

Google already had this. For months.

Here's what barely got covered in the OpenAI announcement. Google didn't just invent SynthID — they already deployed it.

Gemini Nano, Gemini Flash, Gemini Pro, and Gemini Omni all have SynthID embedded at generation time. Google built this, shipped it quietly, and had been running it in production while OpenAI was still figuring out their C2PA compliance.

More importantly — Google built the verification side. You can upload any image to Gemini and ask: “Was this made by Google AI?” Gemini checks both the SynthID watermark in the pixels and the C2PA metadata, and gives you a confidence-weighted answer.

Try it

Try it: Generate any image in Gemini, download it, re-upload to Gemini and ask “Was this image generated by Google AI?” It will detect both the SynthID watermark and C2PA credentials — confirming the image as Gemini-origin. Works across Gemini Nano, Flash, Pro, and Omni.

Google also expanded SynthID detection into Chrome and Google Search — meaning verification is moving to infrastructure level. Every image you see in your browser could soon carry a small provenance indicator, without you needing to visit a verification tool at all.

OpenAI's choice to adopt SynthID — a competitor's technology — is genuinely unusual. It signals that SynthID is robust enough that building a competing watermarking standard would fragment the ecosystem in a way that's bad for everyone. So they collaborated instead.

07The honest part

Where the system still breaks

The announcements won't always say this clearly. Here's the honest breakdown.

Failure 01 · Sophisticated

Adversarial regeneration

Run the image through a custom autoencoder to regenerate it in a new latent space. Washes out the SynthID embedding. Requires real technical skill but works.
Failure 02 · Common

Aggressive manipulation

80% crop + 5 rounds of JPEG compression degrades signal below 40% threshold. “No watermark found” — not “not AI-generated.” Correct but unhelpful.
Failure 03 · Critical gap

Open-source models

Midjourney, Stable Diffusion, ComfyUI, Flux — no SynthID. The most dangerous AI-generated content is generated by tools completely outside this system.
Failure 04 · By design

The “no signal” problem

If no watermark is detected, the tool says “no supported signals found” — not “this is real.” A clean result could mean real, stripped, or from a non-participating generator.
The deeper problem: provenance is an ecosystem problem, not a technology problem.

SynthID is technically impressive. C2PA is architecturally sound. But both require buy-in from every link in the chain — generators, editors, platforms, browsers, end users. The image that goes ChatGPT → downloaded → Photoshop → Telegram → Twitter? It arrives clean. No manifest. Watermark degraded. No definitive claim.

08So what

What this means if you're building

If you're building anything in the AI space — or thinking about where the internet is heading — here's what actually matters.

  1. C2PA is becoming table stakes. The EU AI Act mandates disclosure of AI-generated content. If you're building any image generation tool, C2PA compliance is the path. OpenAI just did it. Adobe's been there. Expect every major tool to follow.
  2. SynthID-equivalent watermarking will become standard. Expect Anthropic, Midjourney, and others to adopt pixel-level watermarking. Open-source models will lag — and that gap is where the most dangerous content will continue to come from.
  3. Verification moves into infrastructure. Google putting SynthID detection in Chrome and Search is the signal. This will become a browser-level feature. Every image, a small provenance indicator — without users needing to know what C2PA or SynthID even mean.
  4. “Made by a human” becomes a differentiator. As AI images flood the internet, cryptographic proof from a C2PA-signed camera — Leica, Nikon, Sony — becomes a commercial differentiator. Photojournalism, commercial photography, stock images: “human-made” will carry a premium.
09Your move

Build with AI. Understand what you're building.

Deep technical breakdowns, build walkthroughs, and real AI product tutorials — for founders, operators, and developers who want to stay on the right side of the gap.

Haroon K M · @ closefuture

Visit Haroon K M →

(Your turn)

When every image can be generated in seconds, what will you accept as proof that one is real?

  • #AI
  • #C2PA
  • #SynthID
  • #ContentProvenance
  • #OpenAI
  • #GoogleDeepMind
  • #AIWatermarking
Next article Build AI for $20

(Contents) · 9 sections · 12 min