Posted by Retr0id 10 hours ago
You will not build a perfect system, or even something near perfect. The best you're going to do is make it so that it's hard to casually present AI photos as real, leaving only the cases where it really matters. In the "best" case, you've just made the public more trusting of photos in general, so that when there's actual money or power on the line that makes jumping through the hoops to fake authenticity worth it, the public is more susceptible.
The best outcome at this point is for everyone to get on the same page that photos have roughly the same probative value now as drawings. Poorly thought out snake oil efforts to prove authenticity are only going to delay that.
People are concerned that the technology will lend additional credence to the last 0.1%. But anyone who thinks about the technology for 2 minutes will see you can just point the camera at the screen. In cases where it really matters (a court of law, internet arguments between nerds) people will know it's not 100% reliable. Locks can be picked, and signatures can be forged, but that doesn't make them useless.
"C2PA Cameras Do Not Survive Contact With Reality" does not survive contact with reality where very, very few users would even think of rooting their phone so they can create signed fake images.
When a motivated malicious user (who doesn't actually need that much resources) will be able to convince people something is authentic because the verification passes when it shouldn't since naive users are primed to believe it by default.
Should we also abolish Pangram, because it's not 100% accurate? Someone might be convinced a text is not AI-generated when it actually is! We should get rid of it rather than fool people into thinking it can be determined accurately. What about antivirus? We should abolish it as well rather than fool people into thinking that their software is ever 100% safe. What about HTTPS? We shouldn't call it "secure" shell because the computer you're connecting to could be compromised! I could go on and on and on.
The median instance of AI image generation isn't evidence in a court case. It's cyberbullying, or deepfakes, or fake news. It's called "slop" because there's a lot of it being churned out at low effort.
While it is always an individual tragedy when people treat each other badly (e.g. through deepfakes and all), the real threat does not exist on that level.
This is about misinformation and disinformation, so we're talking state actors. And with that, the 99.9% hypothesis does not hold true.
Malicious users don't need to root their own phones. They just need to go to fakemyimage dot com, and someone else's rooted phone in a clickfarm-type setup signs it for them. I am not operating such a service myself because I thought it was unnecessary in making my point, but perhaps I will have to reconsider.
Make it a legal requirement to mark AI generated photos and enforce penalties for posting unmarked AI generations. Social media should also mark the country of origin for each post, with the knowledge that posts from your own country are covered by these laws.
I say "nearly", because as soon as you need to keep a private key secure from someone with direct physical access to the device, you're entering dangerous territory. TTBOMK there's no way to make a "perfect black box", so it becomes an arms race between defensive "obfuscation" and tamper detection mechanisms in the one hand and stealth scanning techniques on the other. But this is already the case for TPMs -- that is, the situation is no worse than for an already widely accepted technology.
Yes, software LPEs are a risk -- as they are in every nontrivial computer system. New ones will appear, and old ones will be closed in time, as TFA acknowledges.
Re hardware attacks: The (neat!) glitch injection attack the author describes in the linked "lighter" page only raises the implementation cost of doing image certification properly. For example, if the camera module presented only an interface that dumped raw RGB or JPEG-encoded data plus a digital signature that used a private key known only to the manufacturer, then all that would be required to verify a "downstream" image would be to keep a copy of those original bytes inside the final (potentially cropped, filtered, AI-ed, etc.) image, in the worst case roughly doubling its size on disk (though certainly more efficient schemes could be designed). Any interested third party could then compare the original and final images by eye and decide for themselves whether or not the subsequent processing materially changed the image's "meaning".
Finally: Does the existence of lock picks or bolt cutters render padlocks pointless today? Does it corrode society by encouraging people to mistakenly believe that anything they put behind a $5 padlock will be safe forever? No, and no.
This is ridiculous.
It seems like there's a big disconnect between what C2PA says it's for and what certain journalists think it's for.
However I question the value-add when e.g. the BBC website is already authenticated by nature of being served over HTTPS, and anyone who redistributes BBC content can and should link back to the source.
> It's always been obvious that one could point a camera at a screen, I don't think anyone involved with C2PA has claimed otherwise.
They haven't claimed otherwise exactly, but some have implied it's a solvable problem. Here's where the "learn more" link goes, for when Youtube annotates a video as having C2PA metadata: https://support.google.com/youtube/answer/15446725 (Google is a C2PA Steering Committee member)
> The metadata that leads to a 'Captured with a camera' disclosure is made by a third party (for example, a camera manufacturer). This means that there is some risk that someone could take a photo of another screen showing synthetic content. Because the other screen shows an image that has been modified, it wouldn't be eligible for the 'Captured with a camera' disclosure. This issue is called 'air-gapping'. Camera manufacturers will continue to develop detection measures to prevent 'air-gapping', but the sophistication of those detection measures may vary in the near term.
Interestingly they do not mention any of the other known limitations. Their phrasing is highly weasel-wordy, but the implication is clearly that they imagine picture-of-screen detection to become robust (somehow) in the medium-to-long term.
The value would be in images reposted to social media where the website an show a badge that says it came from a certain source.
At over a decade old, still prescient as ever: https://www.youtube.com/watch?v=HUEvRyemKSg
I don't think completely solving this sort of problem is even possible.
The analog hole is alive and well:)