Apple doesn't address the modified photo replay situation, where you take a picture of an already edited image.
Photoshop / AI-gen an image -> display on a high-resolution monitor -> photograph the monitor with iPhone 18 Pro -> valid Apple Reference image.
To get valid reference photos, you can go to the actual physical location, put the iPhone/monitor in a cardboard box to block external light, then photograph the monitor. Paint the inside of the box using Vantablack (stopping reflections) and cover the LiDAR projector with tape.
I can't wait to see Apple Verified™ photos of UFOs flying over the Golden Gate Bridge.
Claim 7 in this patent application describes how depth sensors are used as part of an image authentication process, which would make such a workaround more difficult:
The Apple Reference Image feature is here launched on iPhone 18 Pro and iPhone 18 Pro Max that both have built-in LiDAR sensors that could be used for this process.
Apple's current implementation doesn't integrate LiDAR. And LiDAR wouldn't be enough here, it's trivial to block the projector and hide the dot pattern. No dot pattern = iPhone thinks the object is far away, which is what happens in landscape photos.
A better fix is to take photos with all three iPhone cameras simultaneously, ideally as a 2-3s video, and use the parallax/multiple perspectives to extract depth information. The video files (Possibly audio too) could also be included with the verified image as additional verification.
They can also prevent photos if iPhone detects the LiDAR sensor is covered, similar to how Meta does it with their camera glasses.
> it's trivial to block the projector and hide the dot pattern. No dot pattern = iPhone thinks the object is far away, which is what happens in landscape photos.
I've never looked at the LiDAR hardware, but where is the emitter in relation to the receiver. Why would the LiDAR not reflect off of whatever you're blocking it with and return a very short flight meaning it was very close?
I don't think it's an either or - additional data signals that need to correlate to authenticate will increase confidence. You can use multiple other signals to evaluate whether something is truly a landscape photo, and in that case not require a LiDAR capture, but if you are inside and at close range then you could assume that it should be part of scoring the authentication.
Similarly, LiDAR alone will help disqualify cases where someone is just taking a picture of e.g. a landscape target of the Golden Gate, but that it shown on a screen 1 meter away.
> A better fix is to take photos with all three iPhone cameras simultaneously, ideally as a 2-3s video, and use the parallax/multiple perspectives to extract depth information.
I think iPhones already do this (although without taking multiple seconds of video). iOS is capable of generating pretty accurate depth data even on devices with no LiDAR unit.
Right but I’d argue that realistically this feature is going to be most useful when taking photos of things reasonably close by, people especially, rather than landscape photography.
> all three iPhone cameras simultaneously, ideally as a 2-3s video, and use the parallax/multiple perspectives to extract depth information.
I think optics could be used to make each camera see a different image.
A video could show shake, which could be verified against readings from the phone's accelerometer -- but you could just hold it still and claim that it was on a tripod.
Using the cameras to estimate depth information is not really a better solution than using the LiDAR sensor because the cameras are so close together that the effective range for parallax depth estimation and the effective range of the LiDAR sensor are very similar.
TL;DR: What is C2PA in 60 seconds
What: An open technical standard for embedding cryptographically signed provenance data inside digital media files.
Who: Created by a coalition founded by Adobe, Arm, BBC, Intel, Microsoft, and Truepic in February 2021.
How: A C2PA Manifest (also called a Content Credential) travels inside the file and records who made it, when, and what tools were used.
Why: Deepfake incidents surged from 500,000 to 8 million cases between 2023 and 2025. Provenance gives media a verifiable chain of custody.
Hm, that's more like digitally signing an image. I'm thinking about something like e.g. encoding/encrypting depth information from multiple images to create a composite that can't be duplicated without the source images.
Still, that means that either the fake target scene and your screen presenting it would need to be outside of LiDAR sensor bounds, or you'd need to find a way to make the depth sensor data conform with your fake scene, both increasing the difficulty of producing a forgery.
> increasing the difficulty of producing a forgery
The problem with this thinking is twofold:
1) Whether it actually meaningfully increases the difficulty of a forgery remains to be seen. Despite their initial language about discerning real events, we see no details here about what scene information is used.
2) It increases the potential value of a forgery because now your forgery is attested by Apple.
So it either makes it easier to defraud people or more worthwhile to put in the effort to defraud people or both. None of those outcomes are great.
A sufficiently light absorbing material is indistinguishable from infinity to a lidar. Or simply add a mirror that redirects it sideways, those can't be detected either.
Or, probably just some optics, like a slanted $9 IR mirror [1] in front of it, to direct the lidar to the sky/absorption box. Then you can point at the high res HDR TV that's probably already in your living room.
Well, first, that's only expensive if you're poor. The world is absolutely full of people who can easily piss away your entire annual income throwing a house party.
But I really mean that if the lidar barely works outdoors anyway then actually you don't need to be 16 feet away at all.
Anyway, one may presume that they've thought about this.
Thought about it and also are bright enough not to fall for the “if a single person dies wearing a seat belt, we should abandon seat belts because they do no good at all” fallacy.
It’s almost certainly possible to fool v1 of this system, for some images, in some contexts. It would be shocking if the first implementation was completely perfect. But maybe it’s better than nothing?
I think this will depend on how it gets used. I can imagine numerous outcomes where it's in fact worse than nothing (significantly more effective blackmail, for instance).
> It’s almost certainly possible to fool v1 of this system, for some images, in some contexts. It would be shocking if the first implementation was completely perfect.
Knowing Apple, they've been working on and testing Apple Reference Image for years.
It being perfect isn't the issue; it's that random people on the internet who are just learning about this assume Apple's engineers haven't already thought about everything (and more) mentioned in this thread.
> It being perfect isn't the issue; it's that random people on the internet who are just learning about this assume Apple's engineers haven't already thought about everything (and more) mentioned in this thread.
Given how many bugs there are in macOS and how long they have remained there, I (who have been writing iOS apps from the release of the first retina iPod touch until AI got good) functionally agree with such people; at best, I think Apple's engineers haven't actually solved everything (and more) mentioned in this thread, even if every one of these things may have come up in discussions and even reached an official backlog or task list or similar.
While that is not quite my bar of confidence when implementing wide-reaching technologies that have numerous unexplored knock-on effects, I guess the calculus must have been different on Infinite Loop recently.
It also doesn't prevent you from staging an image or anything that's existed since photography was invented. But that's not the problem they're trying to solve.
> Today, powerful, widely available AI tools allow users to easily generate or alter photorealistic images to a degree that was difficult to imagine just a few years ago.
Photoshop has existed for decades and so has fake images. This is a low friction way to attest "this image came from an iPhone sensor and Apple approved it". It will still take the usual image forensics to determine if the scene it depicts is legitimate.
> "But that's not the problem they're trying to solve."
It is the problem that they say they're trying to solve, though. They specifically say "where the essential role of a photograph is to prove that something actually happened".
It fails the reasonable person test to say that the "something" in that phrase refers to the act of taking the photo itself.
Likewise in "distinguish between photographs that depict real events and...".
So, in your view, photography has been fatally flawed since the late 1800’s, and mere mitigation of AI image gen are insufficient if they don’t also solve actors impersonating real people?
This is so stupid. This makes it like, a thousand times harder to fake a photo than it would otherwise be. You pedants imagining a way to fake it doesn't change that.
> This makes it like, a thousand times harder to fake a photo than it would otherwise be.
The problem with this thinking is twofold:
1) Whether it actually meaningfully increases the difficulty of a forgery remains to be seen. Despite their initial language about discerning real events, we see no details here about what scene information is used.
2) It increases the potential value of a forgery because now your forgery is attested by Apple.
So it either makes it easier to defraud people or more worthwhile to put in the effort to defraud people or both. None of those outcomes are great.
> Whether it actually meaningfully increases the difficulty of a forgery remains to be seen.
Then maybe let’s save those criticisms until this is in the hands of knowledgeable people who can actually test? I mean, I’m no fan of the direction Apple has gone under Tim Cook, but all else being equal I’m inclined to give them the benefit of the doubt that they may have thought this through over the time it took to build more than a random person speculating on HN who just read a blog post for the first time.
> (…) we see no details here about what scene information is used.
And you assume that everything in a post is the sum total of how it works?
> It increases the potential value of a forgery
By that token, should we also not be adding forgery deterrents to ID cards and bills? After all, if you can fake the preventive measures, “it increases the potential value of a forgery”.
This has been possible since the beginning of photography and yet I can’t think of a single scenario where people have been tricked by a staged photo. Yet every day hundreds of millions of people are being fooled by AI generated photos.
> I can’t think of a single scenario where people have been tricked by a staged photo
Let me introduce you to Sir Arthur Conan Doyle and the Cottingley Fairies[0].
"Doyle was enthusiastic about the photographs, and interpreted them as clear and visible evidence of supernatural phenomena. [...] the photographs were faked, using cardboard cutouts of fairies copied from a popular children's book of the time"
Sony's analogous solution (https://authenticity.sony.net/camera/en-us/) claims 3d depth information is built in, I'm sure Apple could do the same given at least some iPhone models have LiDAR on the back
This would work for close up shots taken on iPhone, but not landscape shots. The infrared dots the iPhone LiDAR projects are too weak to appear over long distances.
Also the dots can be trivially blocked by putting your finger over the sensor, sometimes improving photo quality. I do this frequently when I want to take a photo through a window. The absence of the dot matrix tells the iPhone to focus on the background far away instead of the windowpane.
I see, yeah good point. Perhaps the lack of reliable depth data also be baked into some signed metadata property. Wouldn't tell you definitively if something were fake, but could be a context clue if a particular photo were dubious I suppose.
What they could do instead is record a video while taking a photo, the subtle movements (at least if handheld) might have enough information to get an approximation of depth (parallax).
> I can't wait to see Apple Verified™ photos of UFOs flying over the Golden Gate Bridge.
While I'm on board with you about the inabsolute security of this (relative to what's typically expected of cryptographic systems), the fact that their 'verified' state requires a live certification and can be revoked means that the sensor responsible for obviously faked images will see those images and that device no longer certified.
It all relies a lot on trust in Apple, and integration with Apple, and relatively unmotivated attackers.
You don't even have to travel to the location, you can just spoof GPS. And of course that will only be needed until some eastern european kid gets bored one weekend and the signing keys magically appear on pastebin.
To be fair Apple of all companies have the best shot at pulling it off. They've been perfecting their hardware security for years for other reasons and this is just another way to take advantage of that work. But yes, if someone breaks it then the trust is gone and it casts doubt on all of the photos that were ever captured using the broken system.
We're talking about a 48MP camera here, so finding a sufficiently "high-resolution monitor" to pull this off is probably more difficult than you expect. Especially because without at least a few times as many pixels as the camera, it's likely that there will be detectable moire patterns in the image. My guess is that a physical print is more fruitful, but it's still a pretty tricky task to get high enough dynamic range and such to truly fool the sensor.
This sort of concern is presumably why Apple says "Using a neural network with hidden weights, PCC computes a confidence score for the photograph." I'm assuming that things like moire-patterns from pointing the camera at a screen would be caught by that check.
It's of course physically possible to fool the sensor, but at some point it becomes cheaper to just build a UFO and fly it over the actual Golden Gate Bridge.
It doesn't need to be bulletproof to be very valuable.
However much effort is required to fake it - it's proof that the image is either legit or that much effort went in. There's TONS of cases where it's plausible for someone to have put in the effort to fake a photo with AI (nearly zero effort required) but not remotely plausible that they set up some elaborate high quality photo of a fake.
It's also much more damning if you get caught faking it. Think of the examples where police have been caught posting altered images on social media. The lame excuse that some intern didn't realize it would do more than just upscale the image won't fly if some elaborate setup was required.
Also I think if you're rich enough to be able to coat the inside of a cardboard box with Vantablack, you've probably much easier ways to get misinformation out to the public than photographing a monitor in a cardboard box...
"As part of developing the secure digital negative, PCC computes a confidence score that assesses whether the image has the physical characteristics expected of raw output from our camera sensors. Before the developed reference image is signed, PCC sends the photo GUID, sensor ID, and this confidence score to a companion service, which records them and updates the running score associated with that sensor."
I wouldn't be surprised if it's also possible to detect the differences between a photo of a real scene and a photo of a monitor or printout displaying a photo of a real scene, given they have the raw sensor output.
It's even easier than that. You just wait for someone else to figure out, some photography professional with fancy equipment and a hacker-y mindset, and you pay them to sign your photos for you.
Once a defeat device (a camera pointed at a screen) is functional, whoever has it, can simply automate a "receive API request, display image on screen, photograph it, return signed image" pipeline. A cheap internet service. I'd WAG a hundred thousand signatures per day per phone, limited by the sensor speed.
Since there's no way for anyone, Apple included, to correlate photo signatures with the device that signed them, it's also true there's no way to stop one device from signing millions in bulk. ("...an outside observer cannot determine whether any pair of reference images were taken by the same device..."; "...avoid even implicit public association between different photos taken by the same sensor...")
It's the same economic asymmetry as DRM vs. movie piracy (as soon as one group defeats a technical challenge, millions instantly benefit, at zero marginal cost). Apple has no chance of winning.
Is this really that big of a flaw in this implementation? I don't think it's worth the additional complexity to address it. (Encoding depth information in some way, trying to detect "flat" surfaces, whatever).
Discerning a camera taken image of an image is typically very very easy. The collors/exposure/etc will all be obviously wrong in ways to a human, even without doing any analysis.
You mean that it is sometimes very easy. But it is also sometimes impossible. You seem to be thinking only of poor quality photos of poor quality prints, but there's no basis for assuming those characteristics.
I think the timestamp attestation of the digital negative puts a real hamper on this, because it puts a bounded time window on the photo as part of the cryptographic chain of evidence.
So if you're the proud owner of "literally the only photo of a ridiculously unusual event in a highly public area" which is bounded to either a plausible 15-30 minute window or a sketchy March2026->Now window, people can do something like "hey, gee, did anyone else see that UFO over the golden gate bridge at 3pm?"
plus, you know, the confidence score from their secret neural network, which has an unknown scoring function.
Photoshop / AI-gen an image -> display on a high-resolution monitor -> photograph the monitor with iPhone 18 Pro -> valid Apple Reference image.
To get valid reference photos, you can go to the actual physical location, put the iPhone/monitor in a cardboard box to block external light, then photograph the monitor. Paint the inside of the box using Vantablack (stopping reflections) and cover the LiDAR projector with tape.
I can't wait to see Apple Verified™ photos of UFOs flying over the Golden Gate Bridge.