DEV Community

Cover image for I stripped metadata from AI images to see if YouTube would still flag them
vyixor
vyixor

Posted on

I stripped metadata from AI images to see if YouTube would still flag them

I've been curious about this for a while. Every AI image generator I've tried writes metadata into the output files. Sometimes a full C2PA manifest, sometimes just a "made with X" tag, sometimes a pixel-level watermark buried in the image. The point is that platforms know what they're looking at.

YouTube started relying on this in a serious way this year. If a video or thumbnail has AI metadata attached, YouTube shows a disclosure under it. That's fine for people who want to be transparent. But the way YouTube determines whether content is AI is largely by reading the metadata on the file.

So I ran a test. I generated a batch of images with a common AI tool, ran them through SORO's EXIF remover (a tool I built for a completely different reason), and uploaded the cleaned files to a few platforms to see what happened.

Round one

The flags never appeared. Files with empty EXIF were treated as normal uploads. YouTube, X, and Instagram all missed them. The metadata was the only signal being used, and without it, the systems had nothing to work with.

That held for about six weeks.

Round two

Then things shifted. Same batch of files, same cleaning process, but now they were getting flagged again. Not every time, but often enough that it was obvious something had changed. The platforms had added a second layer that doesn't rely on metadata.

That second layer does something else entirely. It looks at the image itself:

Frequency domain analysis. AI-generated images have statistical quirks in high-frequency detail that natural photos don't have.

Pixel watermarks. Some generators embed invisible marks that survive re-encoding.

Heuristic detection. Compression artifacts, sharpness patterns, the way the model handles faces and text.

None of these are perfect. All of them have false positive rates. But they exist, and they're getting better fast.

The interesting part

Metadata being strippable isn't news. Anyone who works with image files knows it. What's interesting is what this reveals about how these detection systems are built.

Metadata is a soft signal. It's easy to remove, easy to forge, easy to lose by re-saving a file. Any detection that depends on it is one canvas re-encode away from being useless. The C2PA spec, which is the standard everyone is adopting right now, acknowledges this openly. Metadata can be removed, and once it is, the provenance is gone.

The platforms know. Which is why they're shifting to image-based detection. Which is creating its own problems.

Here's the tension nobody wants to name. The same tool that removes metadata for privacy also removes metadata for detection. There's no version of this where those two use cases can be separated.

I strip metadata from photos I post. Not to hide anything. My phone writes GPS coordinates into every image I take, and I'd rather not publish where I live. That's a normal privacy practice. Every security guide says to do it.

But when I strip metadata from an AI image, it defeats the detection too. Same tool, same action, opposite intent. Both uses are legitimate. Both uses are covered by the same button.

Which brings me to the actual point

AI content isn't inherently harmful. I keep seeing that framing pushed and it's wrong. An AI-generated illustration for a personal project, a language model draft of a contract clause, a synthetic voiceover for a video essay. None of these things are bad. Some of them are genuinely useful and would take a human days to produce.

The problem is volume. The internet is drowning. Reddit is a quarter AI slop, YouTube is worse in some categories, Google results are getting harder to read every month. The issue isn't that AI exists. It's that there's no cheap way to tell the good from the noise.

Detection systems are trying to solve that. But they're solving it with a signal that's easy to remove and easy to defeat, and the same tools people use to protect their privacy are the same tools that defeat it. Which means the problem isn't technical. It's structural, and it's not going away.

I don't have a solution. I have a test, a set of results, and a bunch of thoughts that don't fit anywhere else. If you're building in this space, I'd like to hear what you're seeing, especially if you have real numbers on detection accuracy.

Top comments (4)

Collapse
 
aiden11 profile image
Aiden •

Your round-one result is real, but the mechanism you hang it on isn't, and it's checkable in one page.

YouTube's own help page ("Disclosing use of GenAI content", support.google.com/youtube/answer/14328491) lists three separate things that get an AI label applied automatically: content made with YouTube's GenAI tools, content that contains C2PA metadata, and content "that our internal systems detect is AI generated or altered." Metadata is one trigger of three, and the third is an image-side detector the page lists on its own.

So "the metadata was the only signal being used" doesn't follow from the flags not appearing. Your batch not tripping the internal detector is the more economical reading.

That puts pressure on round two too. You read it as the platforms adding a second layer, but the doc already lists that layer, so what changed could be the detector, your sample, or both. As written, the post can't tell those apart.

Which is the one thing missing, and it's small: n. How many uploads per round, how many got flagged, per platform. "Not every time, but often enough" is the only line in the piece that isn't a number, and it's carrying the whole conclusion. Put a denominator under it and this gets much harder to wave off. I'd bet round one was a real null and round two is a real signal, but right now the writeup makes them the same size of claim.

If you've got a number in this space headed into a page or a deck, send it. I trace it to the primary source, free.

Collapse
 
vyixor profile image
vyixor •

fair hit on all of it.

the "metadata was the only signal" line was sloppy. should've said "the only signal I could observe from the outside." youtube's page lists internal detection as a separate trigger and I glossed over that. what I saw in round one is consistent with the internal detector just not firing. different claim than the one I wrote.

same with round two. I framed it as a new layer getting added but it could just be the existing detector improving, or my second batch being easier to catch. can't tell them apart from what I have.

the n is where you got me. I don't actually have real numbers. I ran the test informally and wrote it up from impressions, which is exactly what you're pointing at. "not every time but often enough" was hiding that. that's on me.

appreciate you pulling the source instead of just calling it wrong.

Collapse
 
aiden11 profile image
Aiden •

Straight answer. Rarer than the catch.

You can still get the n. One file, two copies: metadata intact vs through your own remover. Alternate which goes first, ten uploads each on one platform. Count labeled vs not, and how long the label takes to show. That's a rate, and it's the same design that pulls round one apart from round two.

Run it, send me the raw counts before you write it up. I check the claim against the design, free.

Collapse
 
vyixor profile image
vyixor •

For anyone who wants the tool, the EXIF remover is at soroflix.xyz. Happy to answer questions about how it works.

Some comments have been hidden by the post's author - find out more